Fundus image analysis and automatic prediction method and device for ophthalmic diseases based on deep learning and readable storage medium thereof

By optimizing fundus image preprocessing and improving the SwinTransformer model, the problems of uneven fundus image quality and differences in left and right eye features were solved, high-precision automatic prediction of ophthalmic diseases was achieved, and diagnostic accuracy and model generalization capabilities were improved, making it suitable for scenarios with insufficient primary medical resources.

CN120563530BActive Publication Date: 2025-09-26CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511082375.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-26
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

The quality of fundus images varies greatly and traditional preprocessing is limited. The shared labels of the left and right eyes limit the model's learning of the differences in binocular features, resulting in insufficient diagnostic accuracy and generalization ability.

Method used

The fundus image quality is optimized through preprocessing methods such as black edge cropping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering and brightness normalization. The left and right eye labels are separated and reorganized, and an improved model based on SwinTransformer is used to achieve bidirectional interaction of left and right eye features using the cross-fusion attention mechanism. Disease classification is performed in combination with the classification head of the B-spline layer.

Benefits of technology

It significantly improves diagnostic accuracy, enhances the generalization ability of the model, reduces the workload of doctors, and is suitable for scenarios with insufficient primary medical resources, providing fast and reliable auxiliary diagnostic support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563530B_ABST
    Figure CN120563530B_ABST
Patent Text Reader

Abstract

The present invention proposes a method, device, and readable storage medium for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning, which belongs to the field of medical image processing and intelligent diagnosis of ophthalmic diseases. The method obtains the ODIR‑5K binocular fundus image dataset and optimizes the image quality through five pre-processing steps, such as black edge cropping and local contrast enhancement; separates and recombines the left and right eye labels to enhance data diversity; and uses an improved SwinTransformer model containing a cross-fusion attention mechanism and a B‑spline classification head to achieve high-precision prediction of normal and seven common ophthalmic diseases, and evaluates performance through indicators such as accuracy. The present invention effectively solves the problems of poor image quality, insufficient label sharing, and feature fusion, improves diagnostic accuracy, assists clinical decision-making, and reduces the burden on doctors, thus having important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing and intelligent diagnosis of ophthalmic diseases, and in particular to a method, device and readable storage medium thereof for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning. Background Art

[0002] Fundus examination is the most direct and effective means of diagnosing ophthalmic diseases. Through fundus imaging, pathological changes (such as hemorrhage, edema, and arteriosclerosis) in key structures such as the retina, optic nerve, and macular area can be clearly observed, providing an important basis for early diagnosis. However, fundus examination faces many challenges in practical application:

[0003] On the one hand, the quality of fundus images is easily affected by lighting, shooting angle, and the patient's eye condition. There are problems such as uneven brightness, noise interference, and inconsistent binocular image alignment. Traditional preprocessing methods are difficult to fully improve image quality, resulting in insufficient mining of key diagnostic information.

[0004] On the other hand, even in large hospitals, doctors find it difficult to carefully interpret each image due to the huge workload, which can easily lead to missed diagnoses and misdiagnoses. At the same time, fatigue and subjective biases in the human visual system further reduce the reliability and consistency of diagnosis.

[0005] Therefore, developing an efficient and accurate intelligent diagnosis system for ophthalmic diseases to solve the above-mentioned technical pain points has become an important demand in the current field of medical image processing and intelligent diagnosis. Summary of the Invention

[0006] The embodiments of the present invention provide a method, device and readable storage medium for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning. These methods address the problems in the prior art where fundus image quality varies greatly, traditional preprocessing has limited effectiveness, the shared labels of the left and right eyes limit the model's learning of binocular feature differences, and traditional models are unable to fully integrate binocular features, resulting in insufficient diagnostic accuracy and generalization ability.

[0007] The core technology of this invention mainly optimizes the quality of fundus images through five-step preprocessing, including black edge cropping and local contrast enhancement, separates and reorganizes left and right eye labels to enhance data diversity, and uses an improved model based on SwinTransformer (including a cross-fusion attention mechanism and a B-spline classification head) to achieve high-precision automatic prediction of eight types of ophthalmic diseases.

[0008] In a first aspect, the present invention provides a method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning, the method comprising the following steps:

[0009] Step 1: Obtain a binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels;

[0010] Step 2: Preprocess the fundus images of the binocular fundus image dataset. The preprocessing includes black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering, and brightness normalization operations in sequence.

[0011] Step 3: Separate the paired left and right eye images based on the disease keywords in the disease label, and reassemble the separated left and right eye images into a new left and right eye image pair through logical operations;

[0012] Step 4: The left and right eye image pairs obtained in step 3 are input into the improved SwinTransformer model for training. The model uses a cross-fusion attention mechanism to achieve bidirectional interaction between left and right eye features, and outputs the classification results through a classification head containing a B-spline layer.

[0013] The classification head including the B-spline layer includes a linear layer, a B-spline layer, and a fully connected layer. The B-spline layer realizes nonlinear fitting of features and disease categories through piecewise polynomial mapping.

[0014] Step 5: Use accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set to predict ophthalmic diseases.

[0015] Furthermore, the black border cropping in step 2 specifically includes: converting the color fundus image into a grayscale image, generating a clipping mask containing 0 and 1, and identifying and extracting a rectangular area containing key fundus features through the mask to remove the black border at the edge of the image.

[0016] Furthermore, the local contrast enhancement in step 2 specifically includes:

[0017] The fundus image is converted to YUV color space, a brightness enhancement weight table is constructed for the brightness channel, and the contrast of different brightness areas is dynamically adjusted based on the weight table. After the processing is completed, the image is converted back to RGB color space.

[0018] Furthermore, the parameters of the contrast-limited adaptive histogram equalization in step 2 are set as follows: the contrast limiting threshold is 2, the image block size is 8×8, and the local details of the fundus image are enhanced by performing contrast limiting processing and interpolation smoothing on the histogram of each block.

[0019] Furthermore, in step 2, the median filter uses a 3×3 sliding window to remove salt and pepper noise in the image and retain edge information; the brightness normalization maps the image brightness value to the range of [0,1] and resizes the image to 224×224×3.

[0020] Furthermore, the rule of the logical operation in step 3 is: if the labels corresponding to the left and right eye images are both normal categories, then the reorganized label is the normal category; if at least one label in the left and right eye images is a disease category, then the reorganized label is the corresponding disease category.

[0021] Furthermore, the cross-fusion attention mechanism in step 4 is specifically as follows: taking the left eye feature as the query and the right eye feature as the key and value, and taking the right eye feature as the query and the left eye feature as the key and value at the same time, achieving deep fusion of the left and right eye features through bidirectional cross-attention calculation.

[0022] In a second aspect, the present invention provides a deep learning-based fundus image analysis and automatic prediction device for ophthalmic diseases, comprising:

[0023] The acquisition module obtains a binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels;

[0024] The preprocessing module preprocesses the fundus images of the binocular fundus image dataset. The preprocessing includes black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering and brightness normalization operations in sequence;

[0025] The recombination module separates the paired left and right eye images according to the disease-specific keywords in the disease label, and recombines the separated left and right eye images into a new left and right eye image pair through logical operations;

[0026] The training module feeds the left and right eye image pairs obtained by the reassembly module into a model improved by SwinTransformer. This model uses a cross-attention mechanism to achieve bidirectional interaction between left and right eye features and outputs classification results through a classification head containing a B-spline layer. The classification head containing the B-spline layer sequentially includes a linear layer, a B-spline layer, and a fully connected layer. The B-spline layer achieves nonlinear fitting of features and disease categories through piecewise polynomial mapping.

[0027] The evaluation module uses accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set and predict ophthalmic diseases.

[0028] Output module, outputs the prediction results.

[0029] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the above-mentioned deep learning-based fundus image analysis and automatic prediction method for ophthalmic diseases.

[0030] In a fourth aspect, the present invention provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process, and the process includes the above-mentioned deep learning-based fundus image analysis and automatic prediction method for ophthalmic diseases.

[0031] The main contributions and innovations of the present invention are as follows:

[0032] 1. Significantly improved diagnostic accuracy: A five-step preprocessing process effectively improves fundus image quality (accuracy increased to 92.31%). Combined with a cross-fusion attention mechanism, the model fully exploits binocular features, achieving a 97.7% accuracy rate for normal and cataract classifications.

[0033] 2. Enhanced model generalization: The left-right eye label separation and recombination strategy solves the label sharing problem in datasets and improves the model's adaptability to different binocular combinations.

[0034] 3. Improve diagnostic efficiency: Reduce the workload of ophthalmologists and provide fast and reliable support for clinical decision-making, especially in scenarios where primary healthcare resources are insufficient;

[0035] 4. Strong technical practicality: The preprocessing steps take into account both efficiency and effectiveness (such as real-time enhancement of LCE and noise suppression of CLAHE), and the model architecture takes into account both accuracy and computational cost, and has important clinical application value.

[0036] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below so that other features, objects, and advantages of the invention are more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0038] Figure 1 is a flowchart of a method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to an embodiment of the present invention;

[0039] Figure 2 This is a general structural diagram of the model according to an embodiment of the present invention;

[0040] Figure 3 is a data preprocessing flow chart according to an embodiment of the present invention;

[0041] Figure 4 2. This is a schematic diagram of random matching of left and right eyes according to an embodiment of the present invention;

[0042] Figure 5is a model architecture diagram according to an embodiment of the present invention;

[0043] Figure 6 is a heat map of a confusion matrix for medical diagnosis classification according to an embodiment of the present invention;

[0044] Figure 7 FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.

[0046] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0047] In the existing technology, the quality of fundus images is uneven and the traditional preprocessing effect is limited. The shared labels of the left and right eyes limit the model's learning of the differences in binocular features, and traditional models find it difficult to fully integrate binocular features, resulting in insufficient diagnostic accuracy and generalization ability.

[0048] Based on this, the present invention solves the problems existing in the prior art based on an improved model of Swin Transformer.

[0049] Example 1

[0050] The present invention aims to propose a method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning. Figure 1 and Figure 2 , the method comprises the following steps:

[0051] Step 1: Obtain the (ODIR-5K) binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels;

[0052] Step 2: Preprocess the fundus images of the binocular fundus image dataset. The preprocessing includes black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering, and brightness normalization operations in sequence.

[0053] In this embodiment, in order to solve the problem of uneven quality of fundus image data when predicting ophthalmic diseases, such as uneven brightness, noise interference, and inconsistent alignment of binocular images, the present invention proposes an advanced image preprocessing method. This method can not only effectively remove noise and enhance image contrast, but also highlight the lesion area in the image through the self-attention mechanism, thereby significantly improving image quality and providing a higher quality data foundation for subsequent model training and diagnostic analysis. The preprocessing flow chart is shown below. Figure 3 The specific steps are as follows:

[0054] 1) First, we cropped unnecessary feature values ​​from the fundus images of the binocular fundus image dataset, removed all unnecessary black borders to reduce the negative impact of black borders, and retained the key feature values ​​of the central fundus. The cropping process first used the OpenCV library to convert the color image into a grayscale image:

[0055]

[0056] in, 、 、 Represents the color image in The red, green, and blue channel pixel values ​​of the position. For pixels representing white, the pixel value is set to 255; and for pixels corresponding to black, the pixel value becomes 0.

[0057] Generates a clipping mask containing the values ​​0 and 1. If the pixel's value is greater than the specified tolerance (default is 6), the mask value will be set to 1 (True); conversely, if the pixel's value is equal to or lower than the tolerance, the mask value will be set to 0 (False):

[0058]

[0059] A rectangular region is then identified containing rows and columns with pixel values ​​of 1, and the identified rectangular region is extracted from the image in RGB format. This step significantly reduces computational requirements, eliminates irrelevant data, and improves the effectiveness of subsequent analysis.

[0060] 2) Secondly, a local contrast enhancement (LCE) method based on brightness adjustment was used to construct a brightness-based weighted lookup table (LUT), also known as a brightness enhancement weight table. This method analyzes the brightness values ​​of the input pixels and dynamically adjusts the image contrast to meet the enhancement requirements of different brightness areas.

[0061] Specifically, this method divides the brightness range into 7 partitions (such as extremely dark, darker, medium, brighter, extremely bright, etc.), and adopts different enhancement strategies for different brightness levels. In the preprocessing stage, the RGB image is converted to the YUV color space, and the brightness channel (Y channel) is processed independently. The weight table is efficiently applied by using a lookup table (LUT) to reduce the performance consumption of loop calculations. After the enhancement of the brightness channel is completed, the enhanced Y channel is merged back into the image, and it is converted from the YUV color space back to RGB format, and finally converted to PIL format to support subsequent processing. Since the brightness enhancement weight table (LUT) is pre-calculated and mapped pixel by pixel, this method has extremely high computational efficiency and is very suitable for real-time image enhancement. In addition, due to its adaptive characteristics, it can effectively enhance dark details without introducing additional noise, while avoiding overexposure problems in highlight areas:

[0062]

[0063] in is the output intensity after local contrast enhancement; is the low-pass filter value; is the adjustable local gain; is the input pixel brightness.

[0064]

[0065] in is the final output intensity; is the output of the global brightness adjustment.

[0066] 3) Then, the fundus image is subjected to contrast-limited adaptive histogram equalization (CLAHE) to enhance local contrast and detail, making the lesion area more visible. Compared to conventional histogram equalization, CLAHE avoids excessive noise enhancement and suppresses noise amplification by setting a contrast-limiting threshold (CL) and block size (BS). In this example, CL is set to 2 and BS is set to (8×8). The following are the calculation steps of the CLAHE algorithm:

[0067] 1. Divide the image into small blocks of size BS×BS.

[0068] 2. For each small block, calculate its histogram H, where Indicates the number of pixels with intensity value i.

[0069] 3. For each small block of histogram, if , then Reduce to , and reallocate the excess pixels to other bins: in It is the number of other bins (the basic unit for counting and classifying pixel intensities in the histogram, and the core carrier for implementing CLAHE "limiting contrast and equalizing details").

[0070] 4. Calculate the cumulative distribution function (CDF) of each patch and use the CDF to adjust the intensity value of each pixel: in is the position in the original image Pixel intensity value; is the adjusted pixel intensity value; is the number of bins of the histogram.

[0071] 5. Interpolate at the boundaries of small blocks to smooth the transition: where α is the interpolation coefficient, usually between 0 and 1.

[0072] 4) Next, a median filter algorithm is applied to remove salt-and-pepper noise, preserve edge information, smooth the image, reduce interference from sudden pixel changes, and enhance visual quality. This embodiment uses a 3×3 sliding window (i.e., each window contains 9 pixels, 3 rows and 3 columns, centered on the target pixel, and the window traverses the entire fundus image pixel by pixel). This removes noise while better preserving image details and edge information.

[0073]

[0074] in, For The neighborhood of the center; It means that after sorting these 9 pixel values, the value in the middle position (the 5th value) is taken as the grayscale value of the processed central pixel g(x,y); f(i,j) represents the grayscale value of all pixels in the 3×3 neighborhood N centered on (x,y) in the original image (that is, the value of the 9 pixels in the window).

[0075] 5) Finally, use brightness normalization and resize the image to 224*224 pixels. This serves as the input data for the subsequent inference model:

[0076]

[0077] in is the brightness value of the original image; is the minimum brightness value in the image; is the maximum brightness value in the image It is the normalized brightness value, ranging from [0,1].

[0078] Step 3: Separate the paired left and right eye images based on the disease keywords in the disease label, and reassemble the separated left and right eye images into a new left and right eye image pair through logical operations;

[0079] In this embodiment, in order to fully utilize this data structure and improve the generalization ability of the model, the present invention proposes a method for separating the left and right eye labels and randomly combining them. Specifically, by identifying the keywords of the left and right eyes, a corresponding table of descriptions in the keywords and labels is compiled, as shown in Table 1:

[0080] Table 1

[0081]

[0082] Separate the left and right eye labels according to Table 1. During training, the corresponding labels are logically combined according to the task type:

[0083] If both images are normal, the final label is normal;

[0084] If one of the images has one of the seven pathological categories, then an "or" operation is performed and the final normal label is changed to 0.

[0085] Before each round of training, the data will be shuffled and re-randomly matched to enhance the model's adaptability to different combinations. Figure 4 Synthetic labels are shown (first column is "Normal" and the remaining columns are various disease categories).

[0086] Step 4: The left and right eye image pairs obtained in Step 3 are input into a model improved by Swin Transformer for training. This model uses a cross-attention mechanism to achieve bidirectional interaction between left and right eye features and outputs classification results through a classification head containing a B-spline layer. The classification head containing the B-spline layer includes a linear layer, a B-spline layer, and a fully connected layer in sequence. The B-spline layer achieves nonlinear fitting of features and disease categories through piecewise polynomial mapping.

[0087] In this embodiment, the model of the present invention uses the Swin Transformer backbone network as its foundation to extract left and right eye features. It then fuses these features using a self-attention mechanism, inputting them into a classification head that combines an MLP with B-Spline to output the final result. This self-attention fusion method can improve the network's ability to focus on different areas of the binocular image (such as the center and edge areas of the left eye image, and the area corresponding to the right eye image). It can also dynamically adjust the attention weights of binocular image features, allowing the model to focus more on key areas of the binocular image. The MLP combined with B-Spline method can further improve classification accuracy.

[0088] like Figure 5As shown in the figure, the core of the proposed model is the Swin Transformer (Swin-S) backbone network, which is used as a parallel dual-channel feature extractor to process the input left and right eye images separately. The Swin Transformer is a hierarchical visual model based on the Transformer architecture. By introducing a shifted window-based self-attention mechanism, it effectively solves the high computational complexity of traditional Transformers when processing high-resolution images.

[0089] The Swin Transformer works as follows:

[0090] 1. Patch Partition & Embedding: The input image is first divided into non-overlapping patches (Patches), which are regarded as the basic units (tokens) in the sequence and converted into one-dimensional embedding vectors by linear projection.

[0091] 2. Hierarchical feature representation: The model consists of multiple stages. By introducing "patch merging" layers between stages, the model can gradually reduce the spatial resolution of the feature map and increase the number of channels, thereby constructing a hierarchical feature pyramid similar to a convolutional neural network (CNN). It effectively captures rich information from low-level details to high-level semantics.

[0092] 3. Sliding Window Self-Attention (W-MSA & SW-MSA): In each Transformer module, the calculation of self-attention is limited to the local window (W-MSA), and cross-window information interaction is achieved through periodic window shifting (SW-MSA).

[0093] After processing by the Swin-S backbone network, the feature maps are obtained, denoted as left_F and right_F, and their dimensions are both R[B,7,7,768], where B is the batch size (BatchSize).

[0094] The key improvement of this invention lies in its cross-attention fusion mechanism. Unlike allowing features to interact internally, the cross-attention mechanism aims to enable directional information querying between two feature streams. This invention also preserves spatial information, extracting features before the last two layers of Swin-S for feature fusion.

[0095] Specifically, before flattening, the features of one feature stream (for example, the left eye) serve as the query (Query, Q) to focus on and extract relevant information from another feature stream (the right eye), with the latter serving as the key (Key, K) and value (Value, V). This process is bidirectional, and its calculation formula is:

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] In this fusion module, the present invention performs the following two operations:

[0103] 1. Left eye focuses on right eye: Use left_F as query (Q), right_F as key (K) and value (V), and calculate the left eye features enhanced by the right eye information .

[0104] 2. The right eye focuses on the left eye: use right_F as the query (Q), left_F as the key (K) and value (V), and calculate the right eye features enhanced by the left eye information .

[0105] Through this bidirectional cross-attention interaction, the model is able to explicitly model the correspondence and differences between binocular images. These two features that have undergone information interaction are effectively combined (concatenated and then reduced in dimension) to generate the final fused feature, whose dimension remains [B, 256], ensuring the deep integration of binocular information.

[0106] In the classification head part, the present invention adopts the refined structure of "Linear→B-spline→Linear":

[0107] First, the fused features are fed into a linear layer for preliminary feature transformation and dimensionality reduction. Then, the activation function is changed to a B-spline layer, whose input is the output of the previous linear layer. The B-spline function, in its piecewise polynomial form, provides a flexible and smooth nonlinear mapping capability. It can more effectively fit the complex decision boundary between features and categories than traditional activation functions. Finally, the output of the B-spline layer is fed into a final linear layer (i.e., a fully connected layer (FC)) to map the features to a preset number of categories, resulting in the final classification result. This model achieves efficient processing and accurate analysis of binocular images by combining a high-performance Swin-S backbone network, a precise cross-attention fusion mechanism, and an innovative B-spline classification head. This architecture not only retains the powerful feature extraction capabilities of the Swin Transformer, but also enhances the modeling of binocular visual cues through cross-attention. The unique classification head design improves the model's final classification performance.

[0108] Step 5: Use accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set to predict ophthalmic diseases.

[0109] In this embodiment, in order to verify the effect of the model, the following tests were performed:

[0110] 1. ODIR image quality improvement: Five methods are used in sequence to improve the uneven brightness and unclear features of ODIR-5K images.

[0111] In Experiment 1, 3,000 data points were carefully selected from the ODIR-5K dataset and randomly divided into a training set (2,700 cases) and a test set (300 cases) in a 9:1 ratio. Seven different processing methods were designed for the training set, resulting in seven groups: the first group (Original) received no processing; the second group (Cropped) performed black border cropping to remove black edges; the third group performed local contrast enhancement (LCE) to enhance local image detail; the fourth group used limited contrast adaptive histogram equalization (CLAHE) to enhance overall image contrast; the fifth group used median filtering (MF) to remove image noise; the sixth group performed brightness normalization (BN) to normalize image brightness; and the seventh group (ALL-5) used the five methods mentioned above in sequence: black border cropping, local contrast enhancement, CLAHE, median filtering, and brightness normalization. These seven data sets were then fed into the classification model constructed in this study for training.

[0112] After model training was complete, seven sets of trained model parameters were extracted and applied to the test set. During testing, the labels output by the model training were paired with the real labels already in the test set. Finally, by calculating three key metrics: accuracy, precision, and recall, the performance of the model under different treatment methods was comprehensively evaluated. The goal was to screen the optimal treatment solution and provide a scientific and reliable optimal choice for subsequent improvements in model performance efficiency and promotion of clinical application. The results are shown in the following table:

[0113] Table 2 Index evaluation of different methods for fundus image ODIR-5k data preprocessing

[0114]

[0115] It can be seen that each method can improve performance. Among them, CLAHE has a significant improvement, but it needs to solve the problems of brightness imbalance and contrast. The present invention retains each step (five-step method group ALL-5). After the present invention operates in sequence, the effect is the best.

[0116] 2.ODIR Left-Eye Recombination Pairing: Separates the originally paired left-eye and right-eye image pairs based on the left-eye keywords given in the training dataset to enhance the model's generalization and classification capabilities.

[0117] In Experiment 2, 3,000 data points were also selected from the ODIR-5K dataset and randomly divided into a training set (2,700 cases) and a test set (300 cases) in a ratio of 9:1. The labels covered normal categories and seven case categories. Given the small number of samples in the dataset and the particularity that the left and right eye images have been paired, the present invention systematically compiled a case-label matching table corresponding to different descriptions in the keywords based on the detailed descriptions of the cases in the two keywords "Left-DiagnosticKeywords" and "Right-Diagnostic Keywords" in the dataset. Through this table, the originally paired left and right eye image labels were separated, and during the training phase, the left eye label and the right eye label were randomly extracted and re-paired according to specific rules to construct a new left and right eye case combination.

[0118] This experiment involved two comparison groups: the original pairing group (original), which retained the original pairing method without any changes; and the re-paired group (re-paired), which used the re-pairing labeling method described above. Subsequently, these two data sets were fed into the classification model constructed in this study for training (both groups used the data preprocessing method used in Experiment 1, the ALL-5 group).

[0119] After model training is complete, two sets of trained model parameters are extracted and applied to the test set. During the testing phase, the labels output by the model training are paired with the true labels of the test set. Finally, by calculating three key indicators: accuracy, precision, and recall, the performance of the model under different processing methods is comprehensively evaluated. The purpose is to verify the effectiveness of the re-pairing label method and provide a scientific and reliable decision-making basis for subsequent optimization of model performance and promotion of clinical application. Access is as follows Table 3:

[0120] Table 3 Index evaluation of left-right eye re-matching method for fundus image ODIR-5k

[0121]

[0122] It can be seen that the left-right eye recombinant pairing of the present invention significantly improves the performance indicators, which shows that when the data set is insufficient, this method can be used to improve the performance of binocular images.

[0123] 3. ODIR image disease classification: Use a deep learning model based on a Transformer variant to automatically classify ODIR-5K image diseases.

[0124] In Experiment 3, 3,000 data points were also selected from the ODIR-5K dataset and randomly divided into a training set (2,700 cases) and a test set (300 cases) in a 9:1 ratio. Labels were assigned to both normal and seven case categories. The optimal methods from Experiments 1 and 2 (using five methods for data preprocessing and left-right eye re-pairing) were applied to different models for comparison.

[0125] After model training was complete, two sets of trained model parameters were extracted and applied to the test set. During the testing phase, the labels output by the model training were paired with the true labels of the test set. Finally, by calculating three key metrics—accuracy, precision, and recall—the performance of the model under different processing methods was comprehensively evaluated. This aimed to verify the effectiveness of this deep learning model approach and provide a scientific and reliable basis for decision-making in subsequent clinical applications. The results are shown in the following table:

[0126] Table 4 Index evaluation under different models

[0127]

[0128] It can be seen that the present invention has the best performance index.

[0129] 4. Systematic evaluation of binocular data preprocessing methods, left and right eye re-pairing, and overall evaluation of the classification model

[0130] In Experiment 4, this invention was validated using the official test set from the A07 topic of the 16th China University Student Service Outsourcing Innovation and Entrepreneurship Competition, as well as the test set from ODIR. Labels included normal and seven case categories. The best-trained model was used to analyze the prediction performance for different categories (disease types). This analysis aims to identify current shortcomings and prepare for future improvements, paving the way for better integration with hospitals and providing assistance to doctors.

[0131] like Figure 6 The heatmaps of the eight categories (Normal, Diabetes, Glaucoma, Hypertension, Cataract, AMD, Myopia, and Other) show how well the model can identify different disease types. The accuracy rates for each category are as follows:

[0132] Normal: 97.7%

[0133] Cataract: 97.7%;

[0134] Diabetes: 88.8%;

[0135] AMD: 90.2%

[0136] Glaucoma: 89.5%;

[0137] Myopia: 91.9%;

[0138] Hypertension: 83.3%;

[0139] Others: 86.1%;

[0140] Normal and Cataract have the darkest colors in the corresponding diagonal areas of the heat map, with accuracy rates reaching 97.7%. This demonstrates that the model performs very well for these two categories, accurately distinguishing cases from other categories. Myopia (91.9%), Age-Related Macular Degeneration (AMD) (90.2%), and Glaucoma (89.5%) perform well, with darker colors in the diagonal areas of the heat map, demonstrating strong recognition capabilities. Diabetes (88.8%) and Other (86.1%) perform moderately well, but still have clinical application value. Hypertensive Retinopathy (83.3%) performs relatively poorly and requires further optimization. Hypertension accounts for a relatively small proportion of samples in the dataset, limiting exposure to the characteristics of these cases during model training, making it difficult to fully learn their unique and stable patterns.

[0141] In contrast, the normal and cataract categories have ample samples, allowing the model to fully learn their characteristic patterns, resulting in excellent recognition results. Although the diabetes category is relatively large, its features are diverse, making it difficult for the model to correctly identify it. The "Other" category covers a wide range of diseases, and the internal lesion features are complex, diverse, and highly variable, equivalent to a "small dataset collection." Faced with such mixed and heterogeneous features, the model struggles to establish a unified and effective discrimination logic, making recognition more challenging. The normal and cataract categories, on the other hand, have relatively simple and typical features, making them easier for the model to learn and identify.

[0142] While the model still has room for improvement in identifying some disease types, it already possesses significant diagnostic value in assisting doctors. For disease types with excellent recognition (normal, cataract, etc.), the model enables rapid and accurate screening, outputting high-confidence predictions. For disease types with good recognition (myopia, AMD, glaucoma, etc.), the model can help doctors initially identify the underlying disease, significantly reducing the time required for basic diagnosis. For example, when the model identifies a cataract or a normal fundus, doctors can highly trust the result and quickly make diagnostic decisions, significantly improving diagnostic efficiency.

[0143] Example 2

[0144] Based on the same concept, the present invention also proposes a deep learning-based fundus image analysis and automatic prediction device for ophthalmic diseases, comprising:

[0145] The acquisition module obtains a binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels;

[0146] The preprocessing module preprocesses the fundus images of the binocular fundus image dataset. The preprocessing includes black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering and brightness normalization operations in sequence;

[0147] The recombination module separates the paired left and right eye images according to the disease-specific keywords in the disease label, and recombines the separated left and right eye images into a new left and right eye image pair through logical operations;

[0148] The training module feeds the left and right eye image pairs obtained by the reassembly module into a model improved by SwinTransformer. This model uses a cross-attention mechanism to achieve bidirectional interaction between left and right eye features and outputs classification results through a classification head containing a B-spline layer. The classification head containing the B-spline layer sequentially includes a linear layer, a B-spline layer, and a fully connected layer. The B-spline layer achieves nonlinear fitting of features and disease categories through piecewise polynomial mapping.

[0149] The evaluation module uses accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set and predict ophthalmic diseases.

[0150] Output module, outputs the prediction results.

[0151] Example 3

[0152] This embodiment also provides an electronic device, referring to Figure 7 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.

[0153] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.

[0154] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0155] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .

[0156] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the deep learning-based fundus image analysis and automatic prediction methods for ophthalmic diseases in the above embodiments.

[0157] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .

[0158] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0159] The input / output device 408 is used to input or output information.

[0160] Example 4

[0161] This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process. The process includes the fundus image analysis and automatic prediction method of ophthalmic diseases based on deep learning according to embodiment one.

[0162] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.

[0163] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0164] Embodiments of the present invention can be implemented by computer software, which is executable by the data processor of the mobile device, such as in the processor entity, or is implemented by hardware, or is implemented by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer executable components configured to perform the embodiment when the program is running. One or more computer executable components can be at least one software code or a part thereof. In addition, at this point, it should be noted that any box of the logic flow in the figure can represent a program step, or interconnected logical circuits, boxes and functions, or a combination of program steps and logical circuits, boxes and functions. The software can be stored in physical media such as memory chips or storage blocks implemented in the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. Physical media is non-transient media.

[0165] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0166] The above embodiments merely illustrate several embodiments of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning, characterized in that: The following steps are involved: Step 1: Obtain a binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels; Step 2: preprocessing the fundus images of the binocular fundus image dataset, wherein the preprocessing includes sequentially performing black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering, and brightness normalization operations; Step 3: Separate the paired left and right eye images according to the disease keywords in the disease label, and reassemble the separated left and right eye images into a new left and right eye image pair through logical operations; Step 4: The left and right eye image pairs obtained in step 3 are input into the improved SwinTransformer model for training. The model uses a cross-fusion attention mechanism to achieve bidirectional interaction between left and right eye features, and outputs the classification results through a classification head containing a B-spline layer. The classification head including the B-spline layer includes a linear layer, a B-spline layer and a fully connected layer in sequence, and the B-spline layer realizes nonlinear fitting of features and disease categories through piecewise polynomial mapping; Step 5: Use accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set to predict ophthalmic diseases.

2. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to claim 1, characterized in that: The black border cropping in step 2 specifically includes: converting the color fundus image into a grayscale image, generating a clipping mask containing 0 and 1, and identifying and extracting the rectangular area containing the key fundus features through the mask to remove the black border at the edge of the image.

3. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to claim 1, characterized in that: The local contrast enhancement in step 2 specifically includes: The fundus image is converted to YUV color space, a brightness enhancement weight table is constructed for the brightness channel, and the contrast of different brightness areas is dynamically adjusted based on the weight table. After the processing is completed, the image is converted back to RGB color space.

4. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to claim 1, characterized in that: The parameters of the contrast-limited adaptive histogram equalization in step 2 are set as follows: the contrast limit threshold is 2, the image block size is 8×8, and the local details of the fundus image are enhanced by performing contrast limit processing and interpolation smoothing on the histogram of each block.

5. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to claim 4, characterized in that: The median filter in step 2 uses a 3×3 sliding window to remove salt and pepper noise in the image and retain edge information; the brightness normalization maps the image brightness value to the range of [0, 1] and adjusts the image size to 224×224×3.

6. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to claim 1, characterized in that: The rule of the logical operation described in step 3 is: if the labels corresponding to the left and right eye images are both normal categories, the reorganized label is the normal category; if at least one label in the left and right eye images is a disease category, the reorganized label is the corresponding disease category.

7. The method for fundus image analysis and automatic prediction of ophthalmic diseases based on deep learning according to any one of claims 1 to 6, characterized in that: The cross-fusion attention mechanism described in step 4 is specifically as follows: using the left eye feature as the query and the right eye feature as the key and value, and using the right eye feature as the query and the left eye feature as the key and value at the same time, achieving deep fusion of the left and right eye features through bidirectional cross-attention calculation.

8. A deep learning-based fundus image analysis and ophthalmic disease automatic prediction device, characterized in that: include: The acquisition module obtains a binocular fundus image dataset, which contains paired left and right eye fundus images and corresponding disease labels; The preprocessing module preprocesses the fundus images of the binocular fundus image dataset. The preprocessing includes black edge clipping, local contrast enhancement, contrast-limited adaptive histogram equalization, median filtering and brightness normalization operations in sequence; The recombination module separates the paired left and right eye images according to the disease-specific keywords in the disease label, and recombines the separated left and right eye images into a new left and right eye image pair through logical operations; The training module feeds the left and right eye image pairs obtained by the reassembly module into a model improved by SwinTransformer. This model uses a cross-attention mechanism to achieve bidirectional interaction between left and right eye features and outputs classification results through a classification head containing a B-spline layer. The classification head containing the B-spline layer sequentially includes a linear layer, a B-spline layer, and a fully connected layer. The B-spline layer achieves nonlinear fitting of features and disease categories through piecewise polynomial mapping. The evaluation module uses accuracy, precision, and recall metrics to evaluate the performance of the trained model on the test set and predict ophthalmic diseases. Output module, outputs the prediction results.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the deep learning-based fundus image analysis and automatic prediction method for ophthalmic diseases according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes the deep learning-based fundus image analysis and automatic prediction method for ophthalmic diseases according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Gaze estimation detection algorithm based on attention crossing and double-path feature fusion network

    CN116563681A

  • Eye fundus image multi-label classification method, system and device and medium

    CN117711057A