Myopia state classification detection model based on eye fundus image, training method and application
By designing a deep learning model combining dynamic convolution, basic module and ECA attention module, the problem of low accuracy of myopia level classification in the existing technology is solved, and more efficient and stable feature extraction and classification effects are achieved.
Patent Information
- Application Number
- CN202411869726.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
When existing deep learning models distinguish subtle differences such as low-grade, moderate, altitude and ultra-high myopia, the accuracy of feature extraction is insufficient, resulting in low classification accuracy.
A myopia state classification detection model based on fundus images was designed, using a combination of dynamic convolution module, basic module, ECA attention module and full connection layer to extract global and local features through multi-level processing to improve feature extraction and classification accuracy.
It significantly improves the extraction effect of refractive grade-related features, improves the accuracy of myopia grade classification, reduces the amount of calculation and parameter, and makes the model more efficient and stable.
Smart Images

Figure CN120014313A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing and artificial intelligence technology, and in particular to an ophthalmic diagnosis and refractive state detection classification model combined with deep learning, as well as a training method and application. Background Art
[0002] The prevalence of myopia is growing rapidly around the world, especially in East Asia, such as China, South Korea, and Japan. The prevalence of myopia among adolescents has exceeded 80%, of which 10% have developed severe myopia. Excessive close-range eye use, bad eye habits, and frequent use of electronic devices are considered to be the main factors for the increasing prevalence of myopia. How to effectively identify and manage refractive errors has become an important research direction in the field of ophthalmology. Myopia not only reduces the quality of life and work efficiency, but also faces more serious risks and is prone to complications such as retinal detachment, glaucoma, cataracts, and macular holes. Therefore, early and rapid detection and intervention of myopia are crucial.
[0003] Traditional optometry equipment and subjective eye chart tests have high requirements on the skills of operators and the cooperation of patients, resulting in certain subjective deviations in the results. Fundus images are a fast and non-invasive means of ophthalmic examination, which makes them suitable for large-scale eye screening and monitoring. Slight changes in fundus structure may be closely related to the degree of myopia. High myopia usually causes an extension of the axial length and irreversible changes in the fundus structure, which provides a reliable reference for the classification of myopia levels. For example, Varadarajan AV et al. first proposed the use of ResNet neural networks to train fundus images to predict refractive errors in 2018, proving that the features in retinal images can accurately predict myopia. GlennLinde et al. developed a low-cost handheld fundus camera based on a deep learning model in 2023 to predict the degree of myopia. The model accuracy is not high, but such portable devices are suitable for resource-poor areas.
[0004] In summary, the myopia level screening method combining fundus images with deep learning has good operability and portability. Although the existing deep learning model has good image feature extraction capabilities, the accuracy of feature extraction is still insufficient when distinguishing subtle differences such as low, moderate, high and very high myopia, resulting in low classification accuracy. Summary of the invention
[0005] The technical problem to be solved by the present invention is how to improve the accuracy of myopia grade classification.
[0006] The present invention solves the above technical problems through the following technical means:
[0007] The myopia status classification detection model based on fundus images, according to the image processing path, the classification detection model includes a dynamic convolution module, a normalization and activation function module, a basic module 1, a first ECA attention module, a basic module 2, a second ECA attention module, a basic module 3, a normalization and activation function, a maximum pooling layer and a fully connected layer, and finally outputs the result;
[0008] The basic module 1 includes three main modules 1 and 1×1 convolution jump connection. The main module 1 sequentially passes through the ReLU6 activation function, the depth-separable convolution, the ReLU6 activation function, the depth-separable convolution and the maximum pooling layer, and outputs the image after repeating three times to extract the global spatial information;
[0009] The basic module 2 includes four main modules 2 densely connected, and the main module 2 is sequentially subjected to ReLU6 activation function, depthwise separable convolution, ReLU6 activation function, depthwise separable convolution, ReLU6 activation function and depthwise separable convolution, and outputs an image after four repetitions. The inter-layer connection method improves feature reuse and information transmission efficiency;
[0010] The basic module 3 includes the main module 1 and a 1×1 convolution jump connection, and then outputs an image after passing through a depthwise separable convolution, a ReLU6 activation function, a depthwise separable convolution, a ReLU6 activation function and a global average pooling layer.
[0011] Furthermore, the image first passes through two dynamic convolution, normalization and activation function modules, and its output is used as the input of basic module 1.
[0012] Furthermore, the first ECA attention module and the second ECA attention module have the same structure, and the calculation process is: the input feature height is H, the width is W, and the number of channels is C. Global average pooling is first performed on each channel, and feature description vectors are generated for different channels. After pooling, weights are calculated for each channel through one-dimensional convolution to represent the attention of different channels.
[0013] The present invention also provides a training method for a myopia status classification detection model based on fundus images, comprising the following steps: first, the image data needs to be enhanced, and the edge detection is performed on the enhanced image data using a Sobel edge detection algorithm to obtain preprocessed image data; the preprocessed image data is normalized and then input into the classification detection model for training.
[0014] Furthermore, image enhancement includes rotating the image at a random angle, flipping the image horizontally or vertically, and scaling the image at a random ratio.
[0015] The present invention also provides a method for classifying myopia states using the above-mentioned fundus image-based myopia state classification detection model.
[0016] The present invention also provides a myopia state classification system based on a myopia state classification detection model of fundus images, comprising:
[0017] Data processing module: used to pre-process image data;
[0018] Prediction module: uses the data preprocessed by the above data processing module as the input of the above-mentioned myopia status classification detection model, and outputs the glasses myopia classification results.
[0019] Furthermore, the myopia state classification detection model follows the image processing path, and the classification detection model includes a dynamic convolution module, a normalization and activation function module, a basic module 1, a first ECA attention module, a basic module 2, a second ECA attention module, a basic module 3, a normalization and activation function, a maximum pooling layer and a fully connected layer, and finally outputs the result;
[0020] The basic module 1 includes three main modules 1 and 1×1 convolution jump connection. The main module 1 sequentially passes through the ReLU6 activation function, the depth-separable convolution, the ReLU6 activation function, the depth-separable convolution and the maximum pooling layer, and outputs the image after repeating three times to extract the global spatial information;
[0021] The basic module 2 includes four main modules 2 densely connected, and the main module 2 is sequentially subjected to ReLU6 activation function, depthwise separable convolution, ReLU6 activation function, depthwise separable convolution, ReLU6 activation function and depthwise separable convolution, and outputs an image after four repetitions. The inter-layer connection method improves feature reuse and information transmission efficiency;
[0022] The basic module 3 includes the main module 1 and a 1×1 convolution jump connection, and then outputs an image after passing through a depthwise separable convolution, a ReLU6 activation function, a depthwise separable convolution, a ReLU6 activation function and a global average pooling layer.
[0023] Furthermore, the image first passes through two dynamic convolution, normalization and activation function modules, and its output is used as the input of basic module 1.
[0024] Furthermore, the first ECA attention module and the second ECA attention module have the same structure, and the calculation process is: the input feature height is H, the width is W, and the number of channels is C. Global average pooling is first performed on each channel, and feature description vectors are generated for different channels. After pooling, weights are calculated for each channel through one-dimensional convolution to represent the attention of different channels.
[0025] The advantages of the present invention are:
[0026] The significant benefit of the present invention is that the improved deep learning network model is designed according to the characteristics of fundus images, and the extraction effect of refractive grade-related features is significantly improved through the optimized feature extraction module and attention mechanism. The multi-level processing method captures the basic features of fundus images at the shallow level, and mines the information related to the refractive grade at the deep level, improves the classification accuracy, reduces the amount of calculation and parameters while ensuring the classification accuracy, and is more efficient and stable than the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The flowchart of the overall experimental method of the embodiment of the present invention is as follows;
[0028] Figure 2 A diagram of the deep learning network structure proposed in an embodiment of the present invention;
[0029] Figure 3 This is a basic structural diagram of the basic module 1 in an embodiment of the present invention;
[0030] Figure 4 This is a basic structural diagram of the ECA attention module in an embodiment of the present invention;
[0031] Figure 5 This is a basic structural diagram of the basic module 2 in an embodiment of the present invention;
[0032] Figure 6 It is a basic structural diagram of the basic module 3 in the embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] Example 1
[0035] This embodiment provides a myopia classification detection model based on fundus images for myopia classification detection. The model structure is as follows: Figure 2 As shown in the figure, the fundus image after enhancement and preprocessing is input into the deep learning network model, first passing through two dynamic convolution, normalization and activation function modules in sequence, and then its output result is input into basic module 1, ECA attention module, basic module 2, ECA attention module, basic module 3, normalization and activation function, maximum pooling layer and full connection layer, and finally outputs the result.
[0036] Further, such as Figure 3As shown, the basic module 1 combines the residual structure and the depth-wise separable convolution, and is composed of three main modules 1 and 1×1 convolution jump connections. The main module 1 sequentially passes through the ReLU6 activation function, the depth-wise separable convolution, the ReLU6 activation function, the depth-wise separable convolution and the maximum pooling layer, and outputs the image after three repetitions. The basic module 1 is used to extract global spatial information.
[0037] Further, such as Figure 4 As shown, the present invention uses the ECA attention module, and the entire calculation process is divided into four steps: global average pooling, one-dimensional convolution to generate channel weights, one-dimensional convolution and activation weighting. The input feature height is H, the width is W, and the number of channels is C. Global average pooling (GAP) is first performed on each channel, and feature description vectors are generated for different channels. After pooling, weights are calculated for each channel through one-dimensional convolution to represent the attention of different channels.
[0038] Further, such as Figure 5 As shown, the basic module 2 is composed of four main modules 2 densely connected. The main module 2 passes through the ReLU6 activation function, depth-wise separable convolution, ReLU6 activation function, depth-wise separable convolution, ReLU6 activation function and depth-wise separable convolution in sequence, and outputs the image after repeating it four times. The inter-layer connection method improves the feature reuse and information transmission efficiency.
[0039] Further, such as Figure 6 As shown in the figure, basic module 3 consists of convolutional layers and pooling layers, which are used to integrate multi-scale global features. This module outputs the image after passing through the depthwise separable convolution, ReLU6 activation function, depthwise separable convolution, ReLU6 activation function and global average pooling layer after Main module 1 and 1×1 convolution jump connection. One-dimensional convolution reduces the dimensionality of input features and controls the complexity of the model. Through the stacking of multiple deep convolutional layers for deep feature extraction, more complex feature relationships such as microstructures and texture features in fundus images can be effectively identified.
[0040] Example 2
[0041] Based on the myopia state classification detection model of Example 1, this embodiment provides a model training method, such as Figure 1 As shown, the image data needs to be enhanced first. Image enhancement is divided into random angle rotation of the image, horizontal or vertical flipping of the image, and random scaling of the image. Further, the enhanced image data is preprocessed, and edge detection is performed using the Sobel edge detection algorithm. Further, the preprocessed image data is normalized to improve the stability and generalization ability of the model training. The image data after the above processing is input into the improved deep learning network model of the present invention, and compared with the existing deep learning method.
[0042] Example 3
[0043] Based on the myopia status classification detection model trained in Example 2, this example applies the model to perform myopia status classification detection.
[0044] Comparing the application of the myopia state classification detection model of this embodiment with the existing learning method, the model performance is improved. In the experimental results, the accuracy of the Convnext network in the existing deep learning model performance is 0.5467, and that of DenseNet121 is 0.7557. The accuracy of the deep learning network model proposed in the present invention is 0.9104. From the above results, it can be seen that the deep learning network model proposed in the present invention has greatly improved the accuracy of myopia classification detection.
[0045] The deep learning network model proposed in this embodiment performs multi-level processing on the input fundus image. First, the input fundus image undergoes two dynamic convolution and activation function operations to achieve image upsampling and feature enhancement. Normalization accelerates the convergence speed of the model. The activation function limits the activation range to reduce the impact of noise and improves the overall efficiency and stability of the network. Secondly, the model adopts a combination of multiple basic modules and ECA attention modules to gradually accumulate and strengthen features at different levels. The basic module extracts basic features such as vascular structure and optic disc contour in the image through multi-channel convolution operations to help identify features of different refractive levels. At the same time, the ECA attention mechanism increases the attention to important feature channels, adaptively allocates the weights of feature channels, and enables the network to focus on key features related to refractive grade classification, further enhancing the feature expression related to refractive grade. Then, the maximum pooling layer is entered to reduce the dimension of the feature map, while retaining key information, reducing the amount of calculation and effectively suppressing overfitting. Finally, after feature extraction and downsampling, the feature map passes through the fully connected layer, which integrates all extracted features and maps them to specific classification results, and finally completes the classification prediction of myopia level.
[0046] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A myopia status classification detection model based on fundus images, characterized in that: According to the image processing path, the classification detection model includes a dynamic convolution module, a normalization and activation function module, a basic module (1), a first ECA attention module, a basic module (2), a second ECA attention module, a basic module (3), a normalization and activation function, a maximum pooling layer and a fully connected layer, and finally outputs the result; The basic module 1 includes three main modules 1 and 1×1 convolution jump connection. The main module 1 sequentially passes through the ReLU6 activation function, the depth-separable convolution, the ReLU6 activation function, the depth-separable convolution and the maximum pooling layer, and outputs the image after repeating three times to extract the global spatial information; The basic module (2) comprises four main modules (2) densely connected, wherein the main modules (2) are sequentially subjected to ReLU6 activation function, depthwise separable convolution, ReLU6 activation function, depthwise separable convolution, ReLU6 activation function and depthwise separable convolution, and the image is output after four repetitions, and the inter-layer connection method improves feature reuse and information transmission efficiency; The basic module (3) includes the main module (1) and a 1×1 convolution jump connection, and then passes through a depthwise separable convolution, a ReLU6 activation function, a depthwise separable convolution, a ReLU6 activation function and a global average pooling layer to output an image.
2. The myopia state classification detection model based on fundus images according to claim 1 is characterized in that: The image first passes through two dynamic convolution, normalization and activation function modules, and its output is used as the input of the basic module (1).
3. The myopia state classification detection model based on fundus images according to claim 1 is characterized in that: The first ECA attention module and the second ECA attention module have the same structure, and the calculation process is: the input feature height is H, the width is W, and the number of channels is C. Global average pooling is first performed on each channel, and feature description vectors are generated for different channels. After pooling, weights are calculated for each channel through one-dimensional convolution to represent the attention of different channels.
4. The method for training a myopia state classification detection model based on fundus images according to any one of claims 1 to 3, characterized in that: The following steps are involved: First, the image data needs to be enhanced, and the edge detection is performed on the enhanced image data using the Sobel edge detection algorithm to obtain the preprocessed image data; The preprocessed image data is normalized and then input into the classification detection model for training.
5. The method for training a myopia state classification detection model based on fundus images according to claim 4, characterized in that: Image augmentation includes rotating images by random angles, flipping images horizontally or vertically, and scaling images by random ratios.
6. Use the myopia state classification detection model based on fundus images described in any one of claims 1 to 3 to perform myopia state classification.
7. A myopia state classification system based on a myopia state classification detection model of fundus images, characterized in that: include: Data processing module: used to pre-process image data; Prediction module: uses the data preprocessed by the above data processing module as the input of the myopia status classification detection model described in any one of claims 1 to 3, and outputs the glasses myopia classification result.
8. The myopia state classification system based on the myopia state classification detection model of fundus image according to claim 7 is characterized in that: The myopia state classification detection model follows the image processing path, and the classification detection model includes a dynamic convolution module, a normalization and activation function module, a basic module (1), a first ECA attention module, a basic module (2), a second ECA attention module, a basic module (3), a normalization and activation function, a maximum pooling layer and a fully connected layer, and finally outputs a result; The basic module 1 includes three main modules 1 and 1×1 convolution jump connection. The main module 1 sequentially passes through the ReLU6 activation function, the depth-separable convolution, the ReLU6 activation function, the depth-separable convolution and the maximum pooling layer, and outputs the image after repeating three times to extract the global spatial information; The basic module (2) comprises four main modules (2) densely connected, wherein the main modules (2) are sequentially subjected to ReLU6 activation function, depthwise separable convolution, ReLU6 activation function, depthwise separable convolution, ReLU6 activation function and depthwise separable convolution, and the image is output after four repetitions, and the inter-layer connection method improves feature reuse and information transmission efficiency; The basic module (3) includes the main module (1) and a 1×1 convolution jump connection, and then passes through a depthwise separable convolution, a ReLU6 activation function, a depthwise separable convolution, a ReLU6 activation function and a global average pooling layer to output an image.
9. The myopia state classification system based on the myopia state classification detection model of fundus image according to claim 8, characterized in that: The image first passes through two dynamic convolution, normalization and activation function modules, and its output is used as the input of the basic module (1).
10. The myopia state classification system based on the myopia state classification detection model of fundus image according to claim 8, characterized in that: The first ECA attention module and the second ECA attention module have the same structure, and the calculation process is: the input feature height is H, the width is W, and the number of channels is C. Global average pooling is first performed on each channel, and feature description vectors are generated for different channels. After pooling, weights are calculated for each channel through one-dimensional convolution to represent the attention of different channels.