Deep learning-based pathological myopia eye bottom image data amplification method

Through improved positioning strategies and VQ-VAE model, combined with GPT generation technology, the problem of lesion area loss in the amplification of fundus image data of pathological myopia is solved, and high-quality synthetic images are generated, which improves the model training effect and supports the diagnosis and treatment of pathological myopia.

CN120339749APending Publication Date: 2025-07-18OUJIANG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397842.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing pathological myopia fundus image data amplification method is prone to destroy or lose the lesion area when generating training data, resulting in poor model training effect.

Method used

The improved positioning strategy and vector quantized variational autoencoder (VQ-VAE) model are adopted, combined with the GPT model, and the lesion area of healthy fundus images are accurately positioned, pasted and reconstructed, and high-quality synthetic images are generated to enhance the diversity of the data set.

Benefits of technology

Effectively retain the details of the lesion area, improve the effect of model training and the quality of data amplification, and support the diagnosis and treatment of pathological myopia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339749A_ABST
    Figure CN120339749A_ABST
Patent Text Reader

Abstract

The invention discloses a pathological myopia eye bottom image data amplification method based on deep learning, and the method comprises the steps: generating high-quality eye bottom image data through an improved positioning strategy and a vector quantization variational auto-encoder (VQ-VAE) model; the method comprises the steps of using a UNet model to position an optic cup, an optic disk and a macular of a healthy fundus image, calculating relative positions between a focus and the optic cup and the optic disk, and accurately pasting a focus area to a corresponding position of the healthy fundus image. A fundus image is coded and decoded through a VQ-VAE model, a high-quality synthetic image is generated, brand new focus features are generated in combination with a GPT model, and the diversity of a data set is enhanced. According to the method, the quality and diversity of pathological myopia eye bottom image data amplification can be effectively improved, and diagnosis and treatment of pathological myopia are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a method for augmenting fundus image data of pathological myopia based on deep learning, which is used to generate high-quality fundus image data to support the diagnosis and treatment of pathological myopia. Background Art

[0002] Pathological Myopia (PM) is a serious ophthalmic disease that may lead to vision loss. At present, deep learning technology has been widely used in fundus image analysis. However, due to the complexity and diversity of the lesion areas in pathological myopia, existing data augmentation methods (such as CutMix) are prone to damage or lose the lesion areas when generating training data, resulting in poor model training effects. Therefore, there is an urgent need for a data augmentation method that can accurately retain the lesion areas and generate high-quality synthetic images. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for augmenting fundus image data of pathological myopia based on deep learning, which generates high-quality fundus image data through an improved positioning strategy and a Vector Quantized Variational Autoencoder (VQ-VAE) model, and improves the effect of model training.

[0004] Technical Solution

[0005] 1. Positioning Strategy:

[0006] Use a pre-trained UNet model to locate the optic cup, optic disc, and macula in healthy fundus images.

[0007] Label the lesion areas in pathological myopia images and calculate the relative positions between the lesions and the optic cup and optic disc.

[0008] According to the relative position information, accurately paste the lesion areas to the corresponding positions in healthy fundus images, and introduce random offsets and rotations to increase data diversity.

[0009] 2. Fundus Image Reconstruction Based on Vector Quantized Variational Autoencoder:

[0010] Use the VQ-VAE model to encode and decode fundus images to generate high-quality synthetic images.

[0011] Improve the generation quality and training efficiency of the VQ-VAE model through an improved decoder design, discriminator introduction, and training strategy optimization.

[0012] 3. Image Generation Combining with GPT Model:

[0013] Input the discrete token sequence generated by the VQ-VAE into the GPT model to generate new lesion features and enhance the diversity of the dataset.

[0014] By splicing the lesion category and location information, guide the GPT model to generate fundus images of specified lesions. Brief Description of the Drawings

[0015] Figure 1 : Schematic diagram of the AE model

[0016] Figure 2 : Structural diagram of the VQ-VAE model

[0017] Figure 3 : Diagram of the decoder model

[0018] Figure 4 : Schematic diagram of the discriminator

[0019] Figure 5 : Schematic diagram of the combination of VQ-VAE and the autoregressive model

[0020] Figure 6 : Schematic diagram of lesion synthesis

[0021] Figure 7 : Schematic diagram of cropping and mixing

[0022] Figure 8 : Comparison chart of reconstruction effects

[0023] Figure 9 : Comparison chart of choroidal neovascularization synthesis

[0024] Figure 10 : Comparison chart of lacquer crack synthesis

[0025] Figure 11 : Comparison chart of macular atrophy synthesis

[0026] Figure 12 : Effect diagram of typical case generation Detailed Implementation Manner

[0027] Example 1: Data augmentation of macular atrophy lesions

[0028] Step 1.1: Labeling and positioning of macular atrophy lesions

[0029] Lesion labeling: Select images containing macular atrophy (C4) from pathological myopia images and use an image labeling tool to accurately label the macular area. Ensure that the boundaries of the macular area are clear during labeling to avoid omission or mislabeling.

[0030] Relative position calculation: Since macular atrophy is directly located in the macular area, there is no need to calculate the relative position with the optic cup and optic disc, and the coordinate values of the macula are directly used for positioning.

[0031] Step 1.2: Cropping and Pasting of Macular Atrophy Lesions

[0032] Lesion Cropping: According to the annotation information, crop the macular atrophy area from the pathological myopia image. When cropping, retain the blood vessels and background information around the macula to ensure the integrity of the lesion.

[0033] Lesion Pasting: Paste the cropped macular atrophy area onto the macular position of the healthy fundus image. When pasting, introduce a random offset (±10%) and a random rotation (±180 degrees) to increase data diversity.

[0034] Step 1.3: Reconstruction and Optimization of the VQ-VAE Model

[0035] Image Reconstruction: Use the VQ-VAE model to encode and decode the pasted image to generate a high-quality synthetic image. Through an improved decoder design, ensure that the details of the macular atrophy area are retained.

[0036] Evaluation Metrics: Calculate the SSIM, PSNR, and FID metrics of the generated image. The SSIM value reaches 0.857, the PSNR value is 28.266, and the FID value is 19.702, indicating that the quality of the generated image is high.

[0037] Corresponding Drawings:

[0038] Figure 7 : Schematic Diagram of Cropping and Mixing (showing the cropping and pasting process of macular atrophy lesions).

[0039] Figure 11 : Comparison Chart of Macular Atrophy Synthesis (showing the comparison between the synthetic image of the macular atrophy lesion and the original image).

[0040] Example 2: Data Augmentation of Choroidal Neovascularization (CNV) Lesions

[0041] Step 2.1: Annotation and Location of CNV Lesions

[0042] Lesion Annotation: Select the images containing choroidal neovascularization (CNV) from the pathological myopia images, and use an image annotation tool to accurately annotate the CNV area. Ensure that the boundary of the CNV area is clear during annotation to avoid omission or mislabeling.

[0043] Relative Position Calculation: Calculate the relative positions (horizontal offset and vertical offset) between the center point of the CNV lesion and the center points of the optic cup and optic disc.

[0044] Step 2.2: Cropping and Pasting of CNV Lesions

[0045] Lesion Cropping: According to the annotation information, crop the CNV region from the pathological myopia image. When cropping, retain the blood vessels and background information around the CNV to ensure the integrity of the lesion.

[0046] Lesion Pasting: Paste the cropped CNV region to the corresponding position of the healthy fundus image. When pasting, according to the relative position information, ensure that the positional relationship between the CNV region and the optic cup and optic disc remains consistent, and introduce random offsets and rotations.

[0047] Step 2.3: Reconstruction and Optimization of the VQ-VAE Model

[0048] Image Reconstruction: Use the VQ-VAE model to encode and decode the pasted image to generate a high-quality synthetic image. Through an improved decoder design, ensure that the details of the CNV region are retained.

[0049] Evaluation Metrics: Calculate the SSIM, PSNR, and FID metrics of the generated image. An SSIM value of 0.866, a PSNR value of 29.493, and an FID value of 19.702 indicate that the quality of the generated image is relatively high.

[0050] Corresponding Drawings:

[0051] Figure 7 : Schematic Diagram of Cropping and Mixing (showing the cropping and pasting process of the CNV lesion).

[0052] Figure 9 : Comparison Chart of Choroidal Neovascularization Synthesis (showing the comparison between the synthetic image of the CNV lesion and the original image).

[0053] Example 3: Data Augmentation of Lacquer Cracks (LC) Lesions

[0054] Step 3.1: Annotation and Location of LC Lesions

[0055] Lesion Annotation: Select the images containing lacquer cracks (LC) from the pathological myopia images, and use an image annotation tool to accurately annotate the LC region. When annotating, ensure that the boundaries of the LC region are clear to avoid omission or mislabeling.

[0056] Calculation of Relative Position: Calculate the relative positions (horizontal offset and vertical offset) between the center point of the LC lesion and the center points of the optic cup and optic disc.

[0057] Step 3.2: Cropping and Pasting of LC Lesions

[0058] Lesion Cropping: According to the annotation information, crop the LC region from the pathological myopia image. When cropping, retain the blood vessels and background information around the LC to ensure the integrity of the lesion.

[0059] Lesion pasting: Paste the cut LC area to the corresponding position of the healthy fundus image. When pasting, according to the relative position information, ensure that the positional relationship between the LC area and the optic cup and optic disc remains consistent, and introduce random offsets and rotations.

[0060] Step 3.3: Reconstruction and optimization of the VQ-VAE model

[0061] Image reconstruction: Use the VQ-VAE model to encode and decode the pasted image to generate a high-quality synthetic image. Through an improved decoder design, ensure that the details of the LC area are preserved.

[0062] Evaluation metrics: Calculate the SSIM, PSNR, and FID metrics of the generated image. The SSIM value reaches 0.855, the PSNR value is 28.676, and the FID value is 19.702, indicating that the quality of the generated image is high.

[0063] Corresponding attached drawings:

[0064] Figure 7 : Schematic diagram of cropping and mixing (showing the cropping and pasting process of LC lesions).

[0065] Figure 10 : Comparison chart of lacquer crack synthesis (showing the comparison between the synthetic image of LC lesions and the original image).

[0066] Example 4: Generation of new lesions combined with the GPT model

[0067] Step 4.1: Combination of the VQ-VAE and GPT models

[0068] Generation of discrete token sequences: Convert the input image into a discrete token sequence through the encoder of the VQ-VAE, and each token corresponds to a codeword in the codebook.

[0069] Lesion information splicing: Splice the category and location information of the lesion before the compressed encoding of the image to guide the GPT model to generate a fundus image of the specified lesion.

[0070] Lesion category encoding: Map the lesion category (such as CNV, LC, C4) to the corresponding token.

[0071] Lesion position encoding: Map the relative coordinates of the lesion in the image to the corresponding token.

[0072] Step 4.2: Training and generation of the GPT model

[0073] Vocabulary Reconstruction: Compress the vocabulary of GPT2 from 50,257 tokens to 2090 tokens, including 2048 VQ-VAE codebook encodings, 32 lesion location encodings, 5 category encodings, and 5 model start and end token encodings.

[0074] Data Augmentation: Flip the original image vertically, horizontally, and rotate it by 90°, and perform image augmentation by combining lesion pasting operations. Re-encode through the VQ-VAE encoder to increase the data volume to 15 times the original scale.

[0075] Sliding Window Segmentation: Segment the tokens of the complete data into 23 training data with 1024 tokens each to ensure the continuity of context information.

[0076] Step 4.3: Image Generation and Evaluation

[0077] Generation Process: Use the lesion name and coordinates as the input to the GPT model to generate a discrete token sequence. Decode the token sequence into a new fundus image through the VQ-VAE decoder.

[0078] Evaluation of Generation Results: Evaluate the quality of the generated images through the FID metric and visualization results. The FID metric of the generated images is 25.746, indicating that the distribution of the generated images is close to that of the real images and the generation quality is high.

[0079] Corresponding Attached Figures:

[0080] Figure 5 : Schematic Diagram of the Combination of VQ-VAE and Autoregressive Model (Showing the Combination Framework of VQ-VAE and GPT Model).

[0081] Figure 12 : Rendering of Generated Typical Cases (Showing Images of Typical Cases Generated by the GPT Model).

[0082] Example 5: Generation of Mixed Multiple Lesions

[0083] Step 5.1: Annotation and Localization of Multiple Lesions

[0084] Lesion Annotation: Select images containing multiple lesions (such as CNV, LC, C4) from pathological myopia images, and use an image annotation tool to accurately annotate each lesion area.

[0085] Calculation of Relative Position: Calculate the relative positions (horizontal offset and vertical offset) between the center points of each lesion and the center points of the optic cup and optic disc.

[0086] Step 5.2: Cropping and Pasting of Multiple Lesions

[0087] Lesion Cropping: According to the annotation information, each lesion area is cropped from the pathological myopia image. During cropping, the blood vessels and background information around the lesion are retained to ensure the integrity of the lesion.

[0088] Lesion Pasting: The multiple cropped lesion areas are pasted onto the corresponding positions of the healthy fundus image. During pasting, according to the relative position information, the positional relationship between each lesion area and the optic cup and optic disc is ensured to be consistent, and random offsets and rotations are introduced.

[0089] Step 5.3: Reconstruction and Optimization of the VQ-VAE Model

[0090] Image Reconstruction: The VQ-VAE model is used to encode and decode the pasted image to generate a high-quality synthetic image. Through an improved decoder design, the details of each lesion area are ensured to be retained.

[0091] Evaluation Metrics: Calculate the SSIM, PSNR, and FID metrics of the generated image. When the SSIM value reaches 0.860, the PSNR value is 28.900, and the FID value is 20.500, it indicates that the quality of the generated image is relatively high.

[0092] Corresponding Drawings:

[0093] Figure 6 : Schematic Diagram of Lesion Synthesis (showing the synthesis process of multiple lesions).

[0094] Figure 8 : Comparison Chart of Reconstruction Effects (showing the comparison between the multiple-lesion synthetic image and the original image).

[0095] Example 6: Verification of the Reconstruction Accuracy of the VQ-VAE Model

[0096] Step 6.1: Reconstruction of the VQ-VAE Model

[0097] Image Input: The original fundus image is input into the VQ-VAE model to generate a reconstructed image.

[0098] Reconstruction Effect: The quality of the reconstructed image is evaluated through the SSIM, PSNR, and CCR metrics. When the SSIM value reaches 0.868, the PSNR value is 29.493, and the CCR value is 99.010%, it indicates that the quality of the reconstructed image is relatively high.

[0099] Step 6.2: Visualization Comparison of the Reconstructed Image

[0100] Visualization Result: Show the original image, the reconstructed image, and the difference heat map to prove that the VQ-VAE model can accurately restore the lesion area during the reconstruction process.

[0101] Difference analysis: The difference heatmap shows that the differences between the reconstructed image and the original image are mainly concentrated in the edge regions of black pixels, and the reduction degree of the lesion regions is relatively high.

[0102] Corresponding attached drawing:

[0103] Figure 8 : Comparison diagram of reconstruction effect (showing the reconstruction effect of the VQ-VAE model).

[0104] Example 7: Verification of the lesion metastasis ability of VQ-VAE

[0105] Step 7.1: Lesion metastasis experiment

[0106] Experimental design: Use the VQ-VAE model to perform operations of encoding extraction and region exchange on two images (Image A and Image B) to generate a fused image Image G.

[0107] Experimental results: The fused image shows obvious advantages in terms of global structure consistency, blood vessel continuity, and lesion feature retention. The FID value is 19.702, indicating that the fused image has no significant difference from the source dataset.

[0108] Step 7.2: Visual comparison of lesion metastasis

[0109] Visualization results: Show the comparison between the synthesized images containing CNV, LC, and C4 lesions and the pasted images, proving the advantages of the VQ-VAE model in lesion metastasis.

[0110] Corresponding attached drawing:

[0111] Figure 6 : Schematic diagram of lesion synthesis (showing the experimental design of lesion metastasis).

[0112] Figure 9 : Comparison diagram of choroidal neovascularization synthesis (showing the metastasis effect of CNV lesions).

[0113] Figure 10 : Comparison diagram of lacquer crack synthesis (showing the metastasis effect of LC lesions).

[0114] Figure 11 : Comparison diagram of macular atrophy synthesis (showing the metastasis effect of C4 lesions).

[0115] Example 8: Application of the AE model

[0116] Step 8.1: Training of the AE model

[0117] Model training: Use a large number of healthy fundus images to train the AE model so that it can learn the structure and features of fundus images.

[0118] Model Application: Apply the AE model to the compression and reconstruction of fundus images to generate a low-dimensional latent representation.

[0119] Corresponding Drawings:

[0120] Figure 1 : Schematic diagram of the AE model (showing the encoder and decoder structures of the AE model).

[0121] Example 9: Improvement of the VQ-VAE Model

[0122] Step 9.1: Improvement of the Decoder

[0123] Cross-Layer Channel Excitation (SLE) Module: Introduce the SLE module. Through cross-layer connection and channel excitation mechanism, enhance the gradient flow of the decoder and improve the quality of the generated images.

[0124] Decoder Structure: Show the improved decoder structure, which includes convolutional layers and upsampling layers with multiple resolutions.

[0125] Corresponding Drawings:

[0126] Figure 3 : Decoder model diagram (showing the improved decoder structure).

[0127] Example 10: Introduction and Optimization of the Discriminator

[0128] Step 10.1: Design of the Self-Supervised Discriminator

[0129] Local Reconstruction and Global Reconstruction: Design a self-supervised discriminator to enhance the feature extraction ability through local reconstruction and global reconstruction tasks.

[0130] Discriminator Structure: Show the structure of the discriminator, which includes convolutional layers and downsampling layers with multiple resolutions.

[0131] Corresponding Drawings:

[0132] Figure 4 : Discriminator schematic diagram (showing the structure of the self-supervised discriminator).

[0133] Example 11: Optimization of the Training Strategy of the VQ-VAE Model

[0134] Step 11.1: Phased Training

[0135] Phase 1: Train the VQ-VAE part using reconstruction loss, codebook loss, and encoder loss.

[0136] Phase 2: Train the discriminator part to optimize the discriminator through adversarial loss and self-supervised reconstruction tasks.

[0137] Phase 3: Joint training to fine-tune the parameters of the entire model.

[0138] Corresponding attached drawing:

[0139] Figure 2 : VQ-VAE model structure diagram (showing the overall structure of the VQ-VAE model).

[0140] Example 12: Training and generation of GPT model

[0141] Step 12.1: Training of GPT model

[0142] Vocabulary reconstruction: Compress the vocabulary of GPT2 to 2090 tokens to reduce the complexity and computational overhead of the model.

[0143] Data augmentation: Flip the original image up and down, left and right, and rotate it by 90°, and perform image augmentation by combining the lesion pasting operation.

[0144] Step 12.2: Image generation

[0145] Generation process: Use the lesion name and coordinates as the input of the GPT model to generate a discrete token sequence. Decode the token sequence into a new fundus image through the VQ-VAE decoder.

[0146] Evaluation of generation results: Evaluate the quality of the generated image through the FID metric and visualization results.

[0147] Corresponding attached drawing:

[0148] Figure 5 : Schematic diagram of the combination of VQ-VAE and autoregressive model (showing the combination framework of VQ-VAE and GPT model).

[0149] Figure 12 : Effect diagram of typical case generation (showing the typical case image generated by the GPT model).

Claims

1. A method for data augmentation of fundus images of pathological myopia based on deep learning, characterized in that, Including the following steps: Using a pre-trained UNet model to localize the optic cup, optic disc, and macula in healthy fundus images; Labeling the lesion areas in pathological myopia images and calculating the relative positions between the lesions and the optic cup and optic disc; According to the relative position information, accurately paste the lesion areas to the corresponding positions in healthy fundus images and introduce random offsets and rotations to increase data diversity.

2. The method according to claim 1, wherein Using the VQ-VAE model to encode and decode fundus images to generate high-quality synthetic images.

3. The method according to claim 2, wherein Through improved decoder design, discriminator introduction, and training strategy optimization, improve the generation quality and training efficiency of the VQ-VAE model.

4. The method according to claim 1, wherein Input the discrete token sequence generated by VQ-VAE into the GPT model to generate new lesion features and enhance the diversity of the dataset.

5. The method according to claim 4, wherein By concatenating the lesion category and position information, guide the GPT model to generate fundus images of specified lesions.