Eye fundus image classification method and device and computer readable storage medium

By constructing a dual-branch fundus image classification model based on MobileViT, combining multi-scale feature extraction and self-attention mechanisms, the accuracy of multiple eye diseases diagnosis in the prior art is solved, and efficient fundus image feature extraction and classification are achieved.

CN119963907APending Publication Date: 2025-05-09TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510042215.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing fundus image classification methods are difficult to accurately identify multiple eye diseases, and ordinary classification networks are difficult to extract global context information and local features of images, resulting in difficulty in accurately identifying.

Method used

Using MobileViT deep learning network as the infrastructure, a fundus image classification model with a dual-branch structure is built, combining multi-scale feature extraction module, self-attention branch and convolutional neural network branch, and fusing global and local features through cross-reconstruction modules.

Benefits of technology

It realizes accurate and automatic classification of multiple characteristics of fundus images, improves classification accuracy, better captures the global and local characteristics of the image, and is suitable for the diagnosis of various eye diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963907A_ABST
    Figure CN119963907A_ABST
Patent Text Reader

Abstract

The invention discloses an eye fundus image classification method and device and a computer readable storage medium, achieves accurate and efficient feature extraction and classification of eye fundus images, and belongs to the technical field of computer-aided medical treatment. The method comprises the steps of 1, data preparation; s11, establishing a fundus image data set and dividing the fundus image data set into training data, test data and unclassified data; 2, constructing a fundus image classification model; the method comprises the following steps: S21, constructing a dual-branch structure eye fundus image classification model by taking a MobileViT deep learning network as a framework, S22, firstly, carrying out multi-scale feature extraction on an image by using a multi-scale feature extraction module, secondly, carrying out feature extraction on a feature map by using dual branches, and finally, carrying out feature fusion on the feature map output by the dual-branch structure by using a cross reconstruction module, and S23, carrying out feature fusion on the feature map output by the dual-branch structure. A feature map output by the cross reconstruction module passes through a Softmax layer to obtain a classification result; 3, training a fundus image classification model; and step 4, applying the eye fundus image classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer-assisted medical technology, and in particular relates to a fundus image classification method, device and computer-readable storage medium. Background Art

[0003] Using deep learning technology for early screening of ophthalmic diseases can not only save doctors time, expand screening coverage, and reduce the pressure caused by local medical resource shortages, but also reduce the interference of image quality and doctors' subjective judgment with the support of big data, helping doctors to give more accurate and reliable results. Fundus photography technology for basic detection of ophthalmic diseases has been popularized in most areas and does not cause too much financial pressure on patients. Therefore, using deep learning technology to assist in screening of ophthalmic diseases has become an initiative with practical application value and far-reaching significance.

[0004] Although deep learning has made breakthroughs in fundus image diagnosis, most existing research focuses on the identification and diagnosis of a single eye disease. The comprehensive diagnosis of multiple eye diseases is still insufficient and does not conform to the actual process of most eye disease diagnosis in clinical practice. Compared with the classification system of a single fundus disease, multi-disease diagnosis can better meet the actual needs of patients and doctors. However, because the differences between ophthalmic disease images are extremely subtle and the image features are similar, it is difficult for ordinary classification networks to accurately extract image features and achieve accurate identification of ophthalmic diseases. Therefore, the subtle differences in fundus images require the use of more advanced feature extraction methods, more detailed feature relationship construction, more sophisticated attention mechanisms, and the enhancement of more feature information.

[0005] Although CNN is widely used in the field of image recognition, its inherent translation invariance and local receptive field make it difficult to understand global context information when processing images, resulting in difficulty in capturing small details and local features of objects. Vision Transformer (ViT) is a method that applies the Transformer model used to process sequence data to the image field. It uses a self-attention mechanism to learn and understand the correlation and dependency between different parts of the input image. Through self-attention calculation, ViT can simultaneously consider all position information in the image to obtain global context information, thereby better capturing key information and local features in the image, and is suitable for fine-grained image classification tasks. Summary of the invention

[0006] The purpose of the present invention is to provide a fundus image classification method, device and computer-readable storage medium to achieve accurate and efficient feature extraction and classification of fundus images.

[0007] The present invention is achieved by adopting the following technical solutions:

[0008] A fundus image classification method,

[0009] Step 1: Data preparation;

[0010] S11, establish fundus image dataset,

[0011] S12, divide the data set into training data, test data and unclassified data,

[0012] Step 2: Construct a fundus image classification model;

[0013] S21, using MobileViT deep learning network as the framework to build a fundus image classification model, which adopts a dual-branch structure.

[0014] S22, first use the multi-scale feature extraction module to extract multi-scale features from the image, then use the self-attention branch and the convolutional neural network branch to extract features from the feature map, and finally use the cross reconstruction module to fuse the feature maps output by the self-attention branch and the convolutional neural network branch.

[0015] S23, the feature map output by the cross reconstruction module is passed through the Softmax layer to obtain the classification result.

[0016] Step 3: training fundus image classification model;

[0017] S31, using the training data in the data set to train the constructed fundus image classification model,

[0018] S32, using the test data in the data set to test the trained fundus image classification model, when the accuracy of the test data of the trained fundus image classification model reaches the standard requirement, the trained fundus image classification model is the trained fundus image classification model.

[0019] Step 4: Apply fundus image classification model

[0020] S41, inputting the unclassified data into a trained fundus image classification model to obtain a fundus image classification result of the unclassified data.

[0021] Further preferably, S11 also includes preprocessing the data set and adjusting it to images of the same size, the acquired open source fundus images include fundus images of different ages and genders, image preprocessing is performed, the length of the shortest side of the original fundus image is used as the height and width of the cropped image, after recording the height and width of the image, the midpoint of the longer side of the image is used as the center, and 0.5 of the length of the shortest side is extended to both sides for cropping, to obtain the cropped image, and adjust it to an image of 224×224 pixels; the classified fundus images in the S12 data set are divided into training data and test data in a ratio of 4:1, and the unclassified fundus images are divided into unclassified data.

[0022] Further preferably, the Softmax layer in S23 converts the feature dimension into the number of classification categories, and then obtains multiple feature classification results of the entire fundus image.

[0023] Further preferably, the multi-scale feature extraction module in S22 adopts a three-branch strategy: the first branch uses a feature pyramid to extract multi-scale features of the image, and adjusts the number of feature map channels through 3×3 convolution and 1×1 convolution; then three cascaded 5×5 pooling layers with a step size of 1 and a padding of 2 are used to extract image features. After each layer of pooling, feature maps with different receptive fields can be obtained. The feature maps extracted by each layer of pooling are spliced, which can make the multi-scale feature extraction model have a larger receptive field without changing the image resolution. It not only extracts multi-scale features of the image, but also effectively solves the problem of feature redundancy; the second branch uses a shortcut branch to enable the network to directly use the information of the previous layer; The feature map input into the multi-scale feature extraction model is only convolved and directly spliced ​​and fused with the multi-scale features obtained by the first branch, so that the network can directly use the feature information of the previous layer, so that the network can capture richer information and improve the generalization ability of the model; the third branch is the attention branch, which is used to obtain the channel attention weight of the feature map of the input multi-scale feature extraction model and re-weight the captured multi-scale features, that is, the feature map of the input multi-scale feature extraction model is globally averaged pooled, and the channel attention weight is obtained through Sigmoid activation. Finally, the obtained attention weight is feature multiplied with the feature map spliced ​​by the first branch and the second branch to obtain the final image multi-scale feature map.

[0024] Further preferably, the specific operation of training the fundus image classification model in step three is as follows: setting the initial parameters and the number of training iterations of the fundus image classification model, using the training data to train the model, the fundus image classification model predicts the test data during the training process, and calculates the classification accuracy value on the test data, and continuously adjusts the parameters of the fundus image classification model during training until the prediction accuracy of the fundus image classification model for the test data reaches the standard requirements. The trained fundus image classification model is the trained fundus image classification model, and the trained fundus image classification model is used to classify the fundus images on the test data to obtain the classification accuracy.

[0025] Further preferably, in S22, a multi-scale feature extraction module is first used to perform multi-scale feature extraction on the image, and then the self-attention branch and the convolutional neural network branch are used to extract features from the feature map, respectively. Specifically, the fundus image input by S221 is first subjected to a 3×3 convolution, and then the multi-scale feature extraction module is used to capture the multi-scale features of the image. Here, the size of the fundus image features is represented as H×W×C, where H, W, and C are the height, width, and dimension of the fundus image feature tensor, respectively.

[0026] Further preferably, the S222 self-attention branch is responsible for extracting the global features of the image, and the convolutional neural network branch is responsible for extracting the local features of the image; the convolutional neural network branch is provided with a total of 5 layers, and 3×3 convolution is used layer by layer in the convolutional neural network branch to extract features from the feature map of the previous layer; the first two layers of the self-attention branch use depthwise separable convolution for feature extraction, and the last three layers use the MobileViT module to extract the global features of the image, and the input features of each layer in this branch come from the splicing and fusion of the features extracted from the previous layer and the features extracted by the convolutional neural network branch at the same level.

[0027] Further preferably, in S22, the cross reconstruction module is finally used to perform feature fusion on the feature maps output by the self-attention branch and the convolutional neural network branch. Specifically, S223 performs feature fusion on the global features obtained by the self-attention branch and the local features obtained by the convolutional neural network branch; first, the local features of the convolutional neural network branch are point-by-point convolved with the global features of the self-attention branch to increase the number of channels of the feature map; secondly, new features are obtained through the cross reconstruction module, the obtained global features and local features are divided from the channel dimension, and then the features are cross-added, and finally, feature splicing is performed and the number of channels is re-adjusted through 3×3 convolution to form new features to obtain output features. In this way, two different information features can be fully integrated and the information flow between each other is strengthened.

[0028] A fundus image classification device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes a fundus image classification method as described above.

[0029] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for classifying fundus images as described above is implemented.

[0030] The beneficial effects of the present invention are:

[0031] 1. Construct a fundus image classification model. The fundus image classification model is based on MobileViT, which extracts multi-scale features of the image and integrates the self-attention branch and the convolutional neural network branch to enable the fundus image classification model to extract the global and local features of the image. It also uses a cross-reconstruction module for feature fusion to achieve automatic and accurate classification of multiple features of fundus images.

[0032] 2. Compared with other classic CNN classification models, the fundus image classification model constructed by the present invention has better classification accuracy.

[0033] 3. The fundus image classification model constructed by the present invention adopts a dual-branch structure. By introducing the self-attention branch, it can effectively enhance the fundus image classification model's capture of the global features of the image, and is assisted by the convolutional neural network branch to enhance the fundus image classification model's control over the image detail information.

[0034] 4. The fundus image classification model constructed by the present invention adopts a cross-reconstruction module to fuse the global features of the self-attention branch with the local features of the convolutional neural network branch.

[0035] 5. The fundus image classification model constructed by the present invention introduces a multi-scale feature extraction module, which can effectively obtain multi-scale features of the image to complete the interaction, thereby improving the classification ability of various fundus images. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 It is a schematic diagram of the process of the present invention;

[0039] Figure 2 This is a schematic diagram of the network structure of the fundus image classification model of the present invention;

[0040] Figure 3 This is a schematic diagram of the network structure of the multi-scale feature extraction module of the present invention;

[0041] Figure 4 This is a schematic diagram of the network structure of the cross-reconstruction module of the present invention. DETAILED DESCRIPTION

[0042] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0043] In the description, it should be noted that the terms "first" and "second" are only used for descriptive purposes and should not be understood as indicating or implying relative importance. It should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be internal communication between two elements. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.

[0045] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0046] Embodiment 1, a fundus image classification method comprises the following steps:

[0047] Step 1: Data preparation.

[0048] S11, establishing a fundus image dataset.

[0049] S1101, preprocessing the data set using a threshold method, obtaining open-source fundus images including fundus images of different ages and genders, performing image preprocessing, extracting irrelevant areas of the fundus images, that is, taking the length of the shortest side of the original fundus image as the height and width of the cropped image, recording the height and width of the image, taking the midpoint of the longer side of the image as the center, extending 0.5 of the length of the shortest side to both sides for cropping, and obtaining the cropped image;

[0050] S1102, adjusting to an image of the same size. In this embodiment, the image is adjusted to an image of 224×224 pixels.

[0051] The specific operations of data preprocessing include data cleaning, image cropping, image resizing and data enhancement.

[0052] Data cleaning: removing fundus images that do not need to be classified. This has little impact on the model classification effect.

[0053] Crop the image. Since most fundus images in the original data have pure black backgrounds, these backgrounds input into the network will cause information redundancy and reduce the model reasoning speed. Therefore, cropping is performed with the central area of ​​the fundus in the image as the center to remove as much excess background as possible.

[0054] When adjusting the image size, since the fundus images in the original dataset may come from different hospitals and different equipment, the image resolutions vary. In order to facilitate input into the network, they are uniformly scaled to 224×224 images.

[0055] Data enhancement. Due to the class imbalance problem in the dataset, especially the number of normal fundus images is about 30 times that of images of fundus diseases related to retina with hypertension, 8 methods are used to enhance the dataset: rotating the original image 90° counterclockwise, flipping the original image horizontally, adjusting the brightness of the original image by a factor of 1.5, enhancing the color of the original image by a factor of 2.0, adding Gaussian noise to the image, adjusting the saturation of the image by a factor of 1.5, performing Gaussian blur on the image, and sharpening the image by a factor of 2.0.

[0056] S12, dividing the data set into training data, test data and unclassified data, wherein the classified fundus images in the data set are divided into training data and test data in a ratio of 4:1, and the unclassified fundus images are divided into unclassified data.

[0057] Step 2: Construct a fundus image classification model;

[0058] S21, a fundus image classification model is constructed based on the MobileViT deep learning network framework, and the fundus image classification model adopts a dual-branch structure.

[0059] This embodiment uses MobileViT as the basic architecture to construct a network model, namely, a fundus image classification model. The fundus image classification model uses the self-attention branch as the trunk to extract the global features of the fundus image, and adopts a dual-branch method to expand the width of the fundus image classification model and enhance the feature extraction capability of the fundus image classification model.

[0060] S22, first use the multi-scale feature extraction module (SPPSE) to extract multi-scale features of the image, then use the self-attention branch and the convolutional neural network branch to extract features from the feature map, and finally use the cross reconstruction module (CRM) to fuse the feature maps output by the self-attention branch and the convolutional neural network branch.

[0061] The multi-scale feature extraction module adopts a three-branch strategy:

[0062] The first branch uses a feature pyramid to extract multi-scale features of the image, and adjusts the number of feature map channels through 3×3 convolution and 1×1 convolution; then three series of 5×5 pooling, stride 1, padding 2 maximum pooling layers are used to extract image features. After each layer of pooling, feature maps with different receptive fields can be obtained. The feature maps extracted by each layer of pooling are spliced, which can make the multi-scale feature extraction model have a larger receptive field without changing the image resolution. It not only extracts multi-scale features of the image, but also effectively solves the problem of feature redundancy.

[0063] Specific operations such as Figure 3 As shown, the input feature is represented as X∈R H×W×C , where H, W, and C are the height, width, and dimension of the input feature tensor respectively. The initial feature is adjusted by 3×3 convolution and 1×1 convolution to obtain the feature f x (Formula 1) and through three series of maximum pooling layers with a pooling kernel of 5×5, a stride of 1, and a padding of 2, the multi-scale features of the image are extracted, and the extracted multi-scale features f mul11 (Formula 2), f mul12 (Formula 3), f mul13 (4) and adjust it through 1×1 convolution and 3×3 convolution to obtain multi-scale features f mul1 (Formula 5).

[0064] f x =Conv 1×1 (Conv 3×3 (Conv 1×1 (X))) (1)

[0065] f mul11 =MaxPool(f x ) (2)

[0066] f mul12 =MaxPool(f mul11 ) (3)

[0067] f mul13 =MaxPool(f mul12 ) (4)

[0068] f mul1 =Conv 1×1 (Conv 3×3 (Cat(f mul11 ,f mul12 ,f mul13 ,f x ))) (5)

[0069] Among them, Conv 1×1 (·) is a 1×1 convolution, Conv 3×3 (·) is a 3×3 convolution, MaxPool(·) is the maximum pooling with a pooling kernel of 5×5, a stride of 1, and a padding of 2, and Cat(·) is feature concatenation.

[0070] The second branch uses a shortcut branch to enable the network to directly use the information of the previous layer; it only performs convolution operations on the feature map of the input multi-scale feature extraction model, and directly splices and fuses it with the multi-scale features obtained through the first branch, so that the network can directly use the feature information of the previous layer, allowing the network to capture richer information and improve the generalization ability of the model.

[0071] Specific operations such as Figure 3 As shown in the figure, the original features are concatenated with the multi-scale features of the image obtained in the first step after only 1×1 convolution to adjust the number of channels, so that the network can directly use the feature information of the previous layer to obtain the multi-scale features f mul2 (Formula 6).

[0072] f mul2 =Cat(f mul1 ,Conv 1×1 (X)) (6)

[0073] The third branch is the attention branch, which is used to obtain the channel attention weights of the feature map of the input multi-scale feature extraction model and re-weight the captured multi-scale features, that is, the feature map of the input multi-scale feature extraction model is globally average pooled and the channel attention weights are obtained through Sigmoid activation.

[0074] Specific operations such as Figure 3 As shown in the figure, the original features are globally averaged pooled to obtain the weight information on the original feature channel, and then normalized by the Sigmoid function, and multiplied with the multi-scale features of the image obtained in the second step to obtain the final output feature map f mul (Formula 7).

[0075] f mul =Sigmoid(GAP(X))×Conv 3×3 (f mul2) (7)

[0076] Among them, Sigmoid(·) is the Sigmoid activation function, and GAP(·) is the global average pooling.

[0077] Finally, the multi-scale feature extraction module multiplies the obtained attention weight with the feature map spliced ​​by the first branch and the second branch to obtain the final image multi-scale feature map.

[0078] This embodiment uses a cross-reconstruction module to fuse the global features and local features extracted by the self-attention branch and the convolutional neural network branch, so that the fundus image classification model can better balance the global features and local features. The fundus image classification model uses a multi-scale feature extraction module to extract multi-scale and deep features of the image.

[0079] like Figure 2 As shown in the figure, the output features of the multi-scale feature extraction module will be input into the dual-branch structure of the self-attention branch and the convolutional neural network branch for further feature extraction. The convolutional neural network branch consists of convolution operations, and the self-attention branch consists of depthwise separable convolution and MobileViT modules. When applied, it is divided into the following steps:

[0080] The fundus image input by S221 first undergoes a 3×3 convolution and then uses a multi-scale feature extraction module to capture the multi-scale features of the image. Here, the size of the fundus image features is represented as H×W×C, where H, W, and C are the height, width, and dimension of the fundus image feature tensor, respectively.

[0081] The S222 self-attention branch is responsible for extracting the global features of the image, and the convolutional neural network branch is responsible for extracting the local features of the image; the convolutional neural network branch has a total of 5 layers, and 3×3 convolution is used layer by layer in the convolutional neural network branch to extract features from the feature map of the previous layer; the first two layers of the self-attention branch use depthwise separable convolution for feature extraction, and the last three layers use the MobileViT module to extract the global features of the image. The input features of each layer in this branch come from the concatenation and fusion of the features extracted from the previous layer and the features extracted from the convolutional neural network branch at the same level.

[0082] S223 performs feature fusion on the global features obtained by the self-attention branch and the local features obtained by the convolutional neural network branch; first, the local features of the convolutional neural network branch are point-by-point convolved with the global features of the self-attention branch to increase the number of channels of the feature map; secondly, new features are obtained through the cross reconstruction module, the obtained global features and local features are divided from the channel dimension, and then the features are cross-added, and finally the features are spliced ​​and the number of channels is re-adjusted through 3×3 convolution to form new features to obtain output features. In this way, two different information features can be fully integrated and the information flow between each other can be strengthened.

[0083] Specifically, we first analyze the local features f of the convolutional neural network branch. p and the global feature f of the self-attention branch a Perform 1×1 convolution respectively to increase the number of channels of the feature map and reorganize the image features. Figure 4 As shown in Figure 1, the global features and local features from different branches are not directly added or concatenated, but new features are obtained through the cross-reconstruction module, and the obtained global features are divided into f a1 、f a2 (Equation 8), local feature f p Divide into f p1 、f p2 (Equation 9), and then add them separately, and finally concatenate them to get the output feature f out (Formula 10).

[0084] f a1 ,f a2 ,=Split(Conv 1×1 (f a )) (8)

[0085] f p1 ,f p2 ,=Split(Conv 1×1 (f p )) (9)

[0086] f out =Conv 3×3 (Cat((f a1 +f p2 ),(f a2 +f p1 ))) (10)

[0087] Among them, Split(·) is the feature segmentation.

[0088] S23, the feature map output by the cross reconstruction module is passed through the Softmax layer to obtain the classification result. The Softmax layer converts the feature dimension into the number of classification categories, and then obtains multiple feature classification results of the entire fundus image.

[0089] Step 3: training fundus image classification model;

[0090] S31, using the training data in the data set to train the constructed fundus image classification model, setting the initial parameters and the number of training iterations of the fundus image classification model, and using the training data to train the model.

[0091] S32, using the test data in the data set to test the trained fundus image classification model, when the accuracy of the test data of the trained fundus image classification model reaches the standard requirement, the trained fundus image classification model is the trained fundus image classification model.

[0092] Specifically, during the training process, the fundus image classification model predicts the test data and calculates the classification accuracy value on the test data. During the training, the parameters of the fundus image classification model are continuously adjusted until the prediction accuracy of the fundus image classification model for the test data reaches the standard requirements. The trained fundus image classification model is the trained fundus image classification model. The trained fundus image classification model is used to classify the fundus images on the test data to obtain the classification accuracy.

[0093] S33, calculating the index value of the classification effect on the test set for the trained fundus image classification model.

[0094] When evaluating the fundus image classification model, it is necessary to judge the algorithm performance based on reasonable evaluation indicators. The indicators for evaluating the classification efficiency of the model include single epoch training time, test time, parameter quantity and GPU memory. The commonly used indicator value for evaluating the classification effect is accuracy (AC). The calculation formula of accuracy AC is shown in formula (11):

[0095]

[0096] TP represents the number of correctly predicted positive samples; FP represents the number of negative samples incorrectly predicted as positive; TN represents the number of correctly predicted negative samples; and FN represents the number of positive samples incorrectly predicted as negative.

[0097] Step 4: Apply fundus image classification model

[0098] S41, inputting the unclassified data into a trained fundus image classification model to obtain a fundus image classification result of the unclassified data.

[0099] Embodiment 2, a fundus image classification device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes a fundus image classification method as described above.

[0100] Embodiment 3, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the fundus image classification method described above is implemented.

[0101] The above is only a specific implementation of the present invention, which enables those skilled in the art to understand or implement the present invention. Although detailed descriptions are given with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.

Claims

1. A fundus image classification method, characterized in that: Step 1: Data preparation; S11, establish fundus image dataset, S12, divide the data set into training data, test data and unclassified data, Step 2: Construct a fundus image classification model; S21, using MobileViT deep learning network as the framework to build a fundus image classification model, which adopts a dual-branch structure. S22, first use the multi-scale feature extraction module to extract multi-scale features from the image, then use the self-attention branch and the convolutional neural network branch to extract features from the feature map, and finally use the cross reconstruction module to fuse the feature maps output by the self-attention branch and the convolutional neural network branch. S23, the feature map output by the cross reconstruction module is passed through the Softmax layer to obtain the classification result. Step 3: training fundus image classification model; S31, using the training data in the data set to train the constructed fundus image classification model, S32, using the test data in the data set to test the trained fundus image classification model, when the accuracy of the trained fundus image classification model meets the standard requirement, the trained fundus image classification model is the trained fundus image classification model; Step 4: Apply the fundus image classification model; S41, inputting the unclassified data into a trained fundus image classification model to obtain a fundus image classification result of the unclassified data.

2. A fundus image classification method according to claim 1, characterized in that: S11 also includes preprocessing the data set and adjusting it to images of the same size. The open source fundus images obtained include fundus images of different ages and genders. Image preprocessing is performed, and the length of the shortest side of the original fundus image is used as the height and width of the cropped image. After recording the height and width of the image, the midpoint of the longer side of the image is used as the center, and 0.5 of the length of the shortest side is extended to both sides for cropping to obtain the cropped image, and it is adjusted to an image of 224×224 pixels; the classified fundus images in the S12 data set are divided into training data and test data in a ratio of 4:1, and the unclassified fundus images are divided into unclassified data.

3. A fundus image classification method according to claim 1, characterized in that: The Softmax layer in S23 converts the feature dimension into the number of classification categories, and then obtains multiple feature classification results of the entire fundus image.

4. A fundus image classification method according to claim 1, characterized in that: The multi-scale feature extraction module in S22 adopts a three-branch strategy: the first branch uses a feature pyramid to extract multi-scale features of the image, and adjusts the number of feature map channels through 3×3 convolution and 1×1 convolution; then three cascaded 5×5 pooling layers with a step size of 1 and a padding of 2 are used to extract image features. After each layer of pooling, feature maps with different receptive fields can be obtained. The feature maps extracted by each layer of pooling are spliced, which can make the multi-scale feature extraction model have a larger receptive field without changing the image resolution. It not only extracts multi-scale features of the image, but also effectively solves the problem of feature redundancy; the second branch uses a shortcut branch to enable the network to directly use the information of the previous layer; for the input multi-scale The feature map of the feature extraction model is only convolved and directly concatenated with the multi-scale features obtained by the first branch, so that the network can directly use the feature information of the previous layer, so that the network can capture richer information and improve the generalization ability of the model; the third branch is the attention branch, which is used to obtain the channel attention weight of the feature map of the input multi-scale feature extraction model and re-weight the captured multi-scale features, that is, the feature map of the input multi-scale feature extraction model is globally averaged pooled, and the channel attention weight is obtained through Sigmoid activation. Finally, the obtained attention weight is feature multiplied with the feature map concatenated by the first branch and the second branch to obtain the final image multi-scale feature map.

5. A fundus image classification method according to claim 1, characterized in that: The specific operations for training the fundus image classification model in step three are as follows: set the initial parameters and training iterations of the fundus image classification model, use the training data to train the model, and during the training process, the fundus image classification model predicts the test data and calculates the classification accuracy value on the test data. During training, the parameters of the fundus image classification model are continuously adjusted until the prediction accuracy of the fundus image classification model for the test data meets the standard requirements. The trained fundus image classification model is the trained fundus image classification model, and the trained fundus image classification model is used to classify the fundus images on the test data to obtain the classification accuracy.

6. A fundus image classification method according to claim 1, characterized in that: In S22, the multi-scale feature extraction module is first used to extract multi-scale features from the image, and then the self-attention branch and the convolutional neural network branch are used to extract features from the feature map. Specifically, the fundus image input by S221 is first subjected to a 3×3 convolution, and then the multi-scale feature extraction module is used to capture the multi-scale features of the image. Here, the size of the fundus image feature is represented as ,in are the height, width and dimension of the fundus image feature tensor respectively.

7. A fundus image classification method according to claim 6, characterized in that: The S222 self-attention branch is responsible for extracting the global features of the image, and the convolutional neural network branch is responsible for extracting the local features of the image; the convolutional neural network branch has a total of 5 layers, and 3×3 convolution is used layer by layer in the convolutional neural network branch to extract features from the feature map of the previous layer; the first two layers of the self-attention branch use depthwise separable convolution for feature extraction, and the last three layers use the MobileViT module to extract the global features of the image. The input features of each layer in this branch come from the concatenation and fusion of the features extracted from the previous layer and the features extracted from the convolutional neural network branch at the same level.

8. A fundus image classification method according to claim 7, characterized in that: In S22, the cross reconstruction module is finally used to fuse the feature maps output by the self-attention branch and the convolutional neural network branch. Specifically, S223 fuses the global features obtained by the self-attention branch with the local features obtained by the convolutional neural network branch; first, the local features of the convolutional neural network branch are point-by-point convolved with the global features of the self-attention branch to increase the number of channels of the feature map; secondly, new features are obtained through the cross reconstruction module, the obtained global features and local features are divided from the channel dimension, and then the features are cross-added, and finally the features are spliced ​​and the number of channels is re-adjusted through 3×3 convolution to form new features to obtain the output features. In this way, the two different information features can be fully integrated and the information flow between each other can be strengthened.

9. A fundus image classification device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes a fundus image classification method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, a fundus image classification method according to any one of claims 1 to 8 is implemented.