An eye fundus image classification method, device, equipment and storage medium

By preprocessing fundus images and using an image classification network model for fine-grained feature attention processing, the problem of low efficiency in manual evaluation is solved, and automated, efficient and accurate classification is achieved.

CN115984613BActive Publication Date: 2026-01-06YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211664362.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2026-01-06
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

In existing technologies, the classification of fundus images relies on manual evaluation, which leads to low efficiency and frequent misclassification, thus reducing the accuracy of classification.

Method used

After preprocessing the fundus images, a pre-trained image classification network model is used to perform attention processing on fine-grained features, thereby achieving automatic classification.

Benefits of technology

It improves the accuracy and efficiency of fundus image classification, reduces human intervention, and outputs more accurate classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984613B_ABST
    Figure CN115984613B_ABST
Patent Text Reader

Abstract

The application discloses an eye fundus image classification method, device and equipment and a storage medium. The method comprises the following steps: acquiring a first eye fundus image to be classified; performing preprocessing on the first eye fundus image to obtain a second eye fundus image after preprocessing; inputting the second eye fundus image into a pre-trained image classification network model to perform image classification; the image classification network model is used for attention processing of fine-grained features of the second eye fundus image, and classification is performed based on the processing result; and the classification result corresponding to the first eye fundus image is determined according to the output of the image classification network model. Through the technical scheme of the embodiment of the application, automatic classification of the eye fundus image can be realized, and the efficiency and accuracy of the eye fundus image classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for classifying fundus images. Background Technology

[0002] With the development of computer technology, the condition of the fundus can be assessed and analyzed by detecting the retina in fundus images.

[0003] Currently, ophthalmologists typically assess and classify fundus images based on their experience, and use the classification results as reference information for judging the condition of the fundus.

[0004] However, this manual classification method is time-consuming and labor-intensive, and there is a possibility of misclassification, which reduces the accuracy of fundus image classification. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and storage medium for classifying fundus images, thereby achieving automatic classification of fundus images and improving the efficiency and accuracy of fundus image classification.

[0006] According to one aspect of the present invention, a fundus image classification method is provided, the method comprising:

[0007] Obtain the first fundus image to be classified;

[0008] The first fundus image is preprocessed to obtain a preprocessed second fundus image;

[0009] The second fundus image is input into a pre-trained image classification network model for image classification. The image classification network model is used to perform fine-grained feature attention processing on the second fundus image and classify it based on the processing results.

[0010] Based on the output of the image classification network model, the classification result corresponding to the first fundus image is determined.

[0011] According to another aspect of the present invention, a fundus image classification device is provided, the device comprising:

[0012] The first image acquisition module is used to acquire the first fundus image to be classified.

[0013] The second image determination module is used to preprocess the first fundus image to obtain a preprocessed second fundus image;

[0014] The image classification module is used to input the second fundus image into a pre-trained image classification network model for image classification. The image classification network model is used to perform fine-grained feature attention processing on the second fundus image and classify it based on the processing results.

[0015] The classification result determination module is used to determine the classification result corresponding to the first fundus image based on the output of the image classification network model.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the fundus image classification method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the fundus image classification method according to any embodiment of the present invention.

[0021] The technical solution of this invention preprocesses the first fundus image to be classified, and then inputs the preprocessed second fundus image into a pre-trained image classification network model for image classification. The image classification network model performs fine-grained feature attention processing on the input second fundus image and classifies it based on the processing results. This can improve the attention of useful information and suppress useless information, so that the image classification network model can automatically output more accurate classification results, improve the accuracy of image classification, and achieve automatic image classification without human intervention, thereby improving the efficiency of image classification.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a fundus image classification method provided in Embodiment 1 of the present invention;

[0025] Figure 2 These are example images of four first fundus images involved in Embodiment 1 of the present invention;

[0026] Figure 3 This is an example image comparing the first fundus image before and after preprocessing, according to Embodiment 1 of the present invention;

[0027] Figure 4 This is a flowchart of a fundus image classification method provided in Embodiment 2 of the present invention;

[0028] Figure 5 This is a schematic diagram of an image classification network model according to Embodiment 2 of the present invention;

[0029] Figure 6 This is a schematic diagram of an attention processing sub-model according to Embodiment 2 of the present invention;

[0030] Figure 7 This is a schematic diagram of a feature mapping unit according to Embodiment 2 of the present invention;

[0031] Figure 8 This is a schematic diagram of a fine-grained attention unit according to Embodiment 2 of the present invention;

[0032] Figure 9 This is a schematic diagram of a feature attention layer according to Embodiment 2 of the present invention;

[0033] Figure 10 This is a schematic diagram of the structure of a fundus image classification device according to Embodiment 3 of the present invention;

[0034] Figure 11 This is a schematic diagram of the structure of an electronic device that implements the fundus image classification method of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] Example 1

[0038] Figure 1 This is a flowchart of a fundus image classification method provided in Embodiment 1 of the present invention. This embodiment is applicable to the classification of fundus images. The method can be executed by a fundus image classification device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0039] S110. Obtain the first fundus image to be classified.

[0040] Here, fundus image can refer to a fundus image including the retina. First fundus image can refer to the fundus image to be classified. Figure 2 Four example images of the first fundus are provided. See also Figure 2 Each first fundus image contains the retina.

[0041] S120. Preprocess the first fundus image to obtain the preprocessed second fundus image.

[0042] The second fundus image can be obtained by preprocessing the first fundus image. The second fundus image has significantly higher image contrast than the first fundus image.

[0043] Specifically, the first fundus image is preprocessed by Gaussian filtering to obtain the preprocessed second fundus image, thereby removing noise from the first fundus image.

[0044] S130. Input the second fundus image into the pre-trained image classification network model for image classification.

[0045] The image classification network model is used to perform fine-grained feature attention processing on the second fundus image and classify it based on the processing results. The image classification network model is trained based on fundus image samples and the preset image classification results corresponding to the fundus image samples.

[0046] Specifically, the second fundus image is input into a pre-trained image classification network model for image classification, and the output of the image classification network model is obtained. The image classification network model then performs fine-grained feature attention processing on the input second fundus image, suppresses useless information, and thus improves the accuracy of the output of the image classification network model.

[0047] S140. Based on the output of the image classification network model, determine the classification result corresponding to the first fundus image.

[0048] The classification result can be the classification result of the fundus image. For example, the classification result can be a normal fundus result or an abnormal fundus result. Furthermore, abnormal fundus results can be further classified according to different fundus conditions. Specifically, based on the output of the image classification network model, the classification result corresponding to the first fundus image is determined, and the determined classification result is fed back to the user for reference.

[0049] It should be noted that image classification network models can be used to predict the grade of diabetic retinopathy. Diabetic retinopathy (DR) is an eye disease associated with diabetes. Approximately 40% to 45% of diabetic patients have this disease to varying degrees. Based on disease characteristics such as lesions in the user's fundus images, diabetic retinopathy can be classified into five grades: normal, mild non-proliferative phase, moderate non-proliferative phase, severe non-proliferative phase, and proliferative phase.

[0050] The technical solution of this invention preprocesses the first fundus image to be classified, and then inputs the preprocessed second fundus image into a pre-trained image classification network model for image classification. The image classification network model performs fine-grained feature attention processing on the input second fundus image and classifies it based on the processing results. This can improve the attention of useful information and suppress useless information, so that the image classification network model can automatically output more accurate classification results, improve the accuracy of image classification, and achieve automatic image classification without human intervention, thereby improving the efficiency of image classification.

[0051] Based on the above technical solution, S120 may include: performing Gaussian filtering on the first fundus image to obtain a filtered third fundus image; and determining a second fundus image based on the first fundus image, the third fundus image, and a preset bias pixel value.

[0052] The third fundus image can be obtained by applying a Gaussian filter to the first fundus image. The preset bias pixel value can be a parameter value that is pre-set to balance the pixel bias in the image.

[0053] Specifically, a Gaussian filter is applied to the first fundus image (I(x,y)) based on a Gaussian filter of preset size σ to obtain a filtered third fundus image (G(x,y;σ)*I(x,y)). The third fundus image is then weighted based on feature emphasis parameters. The first fundus image is also weighted based on local average color adjustment parameters. The weighted first fundus image, the weighted third fundus image, and a preset bias pixel value are added together to determine the second fundus image (I'(x,y;σ)), thereby unifying image brightness and increasing image contrast. Where I'(x,y;σ)=αI(x,y)+βG(x,y;σ)*I(x,y)+γ.

[0054] For example, Figure 3 An example image comparing the first fundus image before and after preprocessing is provided. See also... Figure 3 The brightness of each first fundus image varies. After preprocessing the first fundus images, a second fundus image with reduced brightness differences is obtained, thereby improving the image quality of the second fundus image input into the image classification network model and thus improving image contrast.

[0055] Example 2

[0056] Figure 4 This is a flowchart of a fundus image classification method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment describes in detail the process of image classification using an image classification network model. Explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here. Figure 4 As shown, the method includes:

[0057] S210. Obtain the first fundus image to be classified.

[0058] S220. Preprocess the first fundus image to obtain the preprocessed second fundus image.

[0059] S230. Input the second fundus image into the feature extraction sub-model for feature extraction to obtain the extracted first fundus feature map.

[0060] The image classification network model is used to perform fine-grained feature attention processing on the second fundus image and classify it based on the processing results. The image classification network model can include a feature extraction sub-model, an attention processing sub-model, and a classification processing sub-model. Figure 5 A schematic diagram of an image classification network model is given. The first fundus feature map can be a map containing global features from the second fundus image. The resolution of each second fundus image varies.

[0061] Specifically, see Figure 5 The second fundus image is input into the feature extraction sub-model to extract features from it. For example, the feature extraction sub-model can use ResNet50 as the backbone network. The deepest output of the convolutional layers in the feature extraction sub-model is obtained and determined as the first fundus feature map (F), thereby minimizing the resolution of the second fundus image to unify the image resolution. Where F∈R C×H×W And C represents the number of input channels, H represents the height, and W represents the width of the input feature.

[0062] S240. Input the first fundus feature map into the attention processing sub-model for fine-grained attention processing to obtain the processed second fundus feature map.

[0063] The second fundus feature map can refer to the fundus feature map obtained by performing fine-grained attention processing on the first fundus feature map.

[0064] Specifically, the first fundus feature map is input into the attention processing sub-model, and attention processing is performed on the first fundus feature map based on a preset fine granularity to obtain the processed second fundus feature map. This allows us to focus only on the features in the feature map that are related to the classification, suppress redundant information in the feature map, and thus improve the accuracy of the image classification results.

[0065] For example, S240 may include: inputting the first fundus feature map into the feature mapping unit for feature mapping to obtain the feature-mapped third fundus feature map; inputting the third fundus feature map into the fine-grained attention unit for fine-grained attention processing to obtain the processed second fundus feature map.

[0066] The attention processing sub-model can include a feature mapping unit and a fine-grained attention unit. The third fundus feature map can be the fundus feature map obtained by mapping the features of the first fundus feature map. Figure 6 A schematic diagram of an attention processing sub-model is given.

[0067] Specifically, see Figure 6 The first fundus feature map is input into the feature mapping unit, where global features in the first fundus feature map are feature-mapped to obtain a third fundus feature map with the same number of channels. The third fundus feature map is then input into the fine-grained attention unit, where fine-grained attention processing is performed based on a preset fine-grained setting to obtain a processed second fundus feature map. This allows for focusing on hierarchical-related features in the feature map based on the preset fine-grained setting, suppressing redundant information in the feature map, and further improving the accuracy of image classification results.

[0068] S250. Input the second fundus feature map into the classification processing sub-model for fully connected processing to determine the classification result corresponding to the first fundus image.

[0069] The classification sub-model can be a sub-model constructed from fully connected layers.

[0070] Specifically, the second fundus feature map is input into the classification processing sub-model, the second fundus feature map is weighted, and the weighted result is input into the next connected layer until the unique classification result corresponding to the first fundus image is output. Thus, the first fundus image is classified based on multiple features, thereby improving the accuracy of the image classification result.

[0071] It should be noted that the input to the classification processing sub-model can be the max pooling result of the second fundus feature map after max pooling, or it can be the one-dimensional data corresponding to the max pooling result.

[0072] S260. Based on the output of the image classification network model, determine the classification result corresponding to the first fundus image.

[0073] The technical solution of this invention involves inputting a second fundus image into a feature extraction sub-model for feature extraction to obtain an extracted first fundus feature map, thereby minimizing the resolution of the second fundus image to unify the image resolution. The first fundus feature map is then input into an attention processing sub-model for fine-grained attention processing to obtain a processed second fundus feature map, allowing focus only on hierarchical-related features and suppressing redundant information. Finally, the second fundus feature map is input into a classification processing sub-model for fully connected processing to determine the classification result corresponding to the first fundus image. This allows for image classification processing of the first fundus image based on multiple features, thereby improving the accuracy of the image classification results.

[0074] Based on the above technical solution, "inputting the first fundus feature map into the feature mapping unit for feature mapping to obtain the feature-mapped third fundus feature map" may include: inputting the first fundus feature map into the first average pooling layer for average pooling processing to obtain the average pooled fourth fundus feature map; inputting the fourth fundus feature map into the first multilayer perceptron for processing to obtain the processed fifth fundus feature map; inputting the first fundus feature map into the first max pooling layer for max pooling processing to obtain the max pooled sixth fundus feature map; inputting the sixth fundus feature map into the second multilayer perceptron for processing to obtain the processed seventh fundus feature map; inputting the fifth and seventh fundus feature maps into the activation processing layer for activation processing to obtain the activated eighth fundus feature map; and inputting the first and eighth fundus feature maps into the first residual processing layer for residual processing to obtain the feature-mapped third fundus feature map.

[0075] The feature mapping unit may include a first average pooling layer, a first max pooling layer, a first multilayer perceptron, a second multilayer perceptron, an activation processing layer, and a first residual processing layer. Figure 7 A schematic diagram of a feature mapping unit is given. Both the first average pooling layer and the first max pooling layer can be used for downsampling. The first and second multilayer perceptrons can be used to transform multiple inputs into a single output. The fourth fundus feature map can refer to the downsampled fundus feature map. The fifth fundus feature map can be the feature map output by the first multilayer perceptron. The sixth fundus feature map can refer to the downsampled fundus feature map. The sixth fundus feature map has different prominent features from the fourth fundus feature map. The seventh fundus feature map can be the feature map output by the second multilayer perceptron. The fifth fundus feature map differs from the seventh fundus feature map. The eighth fundus feature map can refer to the feature map obtained by activating two feature maps.

[0076] Specifically, see Figure 7The first fundus feature map (F) is input into the first average pooling layer, and average pooling is performed on the first fundus feature map to obtain the fourth fundus feature map (F). avg (F)); The fourth fundus feature map (F) avg (F) is input to the first multilayer perceptron for processing to obtain the processed fifth fundus feature map (MLP(F)). avg (F)); The first fundus feature map (F) is input into the first max pooling layer for max pooling processing to obtain the sixth fundus feature map (F) after max pooling. max (F)); The sixth fundus feature map (F) max (F) is input to the second multilayer perceptron for processing, resulting in the processed seventh fundus feature map (MLP(F)). max (F))); The fifth fundus feature map (MLP(F) avg (F))) and the seventh fundus feature map (MLP(F) max (F))) is input to the activation processing layer for weighted processing, and the weighted result is normalized based on a preset function, such as the sigmoid function, to achieve activation processing and obtain the activated eighth fundus feature map (σ(MLP(F)). avg (F))+MLP(F max (F))));The first fundus feature map (F) and the eighth fundus feature map (σ(MLP(F)) avg (F))+MLP(F max (F) is input into the first residual processing layer for element-wise multiplication, and the result of the multiplication is used as the third fundus feature map (F') after feature mapping. And F'∈R C×H×W .

[0077] Based on the above technical solution, "inputting the third fundus feature map into a fine-grained attention unit for fine-grained attention processing to obtain a processed second fundus feature map" may include: inputting the third fundus feature map into a feature segmentation layer for feature map segmentation processing to obtain multiple first feature sub-maps in the third fundus feature map; inputting each first feature sub-map into a feature attention layer for feature attention processing to obtain a processed second feature sub-map; and inputting each second feature sub-map into a feature stitching layer for feature sub-map stitching processing to obtain a processed second fundus feature map.

[0078] The fine-grained attention unit can include a feature segmentation layer, a feature attention layer, and a feature concatenation layer. Figure 8A schematic diagram of a fine-grained attention unit is provided. The first feature sub-image can be a feature sub-image obtained by segmenting the third fundus feature image. The second feature sub-image can be a feature sub-image that has undergone feature attention processing. The feature image after stitching the first feature sub-images has the same size as the third fundus feature image. The feature image after stitching the second feature sub-images has the same size as the second fundus feature image.

[0079] Specifically, see Figure 8 The third fundus feature map is input into the feature segmentation layer. Based on a preset number of anchor frames (N), the third fundus feature map is segmented to obtain N first feature sub-maps. The preset number of anchor frames N can be a square of a natural number. For example, the preset number of anchor frames N can be 9, 16, or 25, etc. Taking a preset number of anchor frames N of 16 as an example, the third fundus feature map is segmented to obtain 16 first feature sub-maps. The i-th first feature sub-map can be represented as ABi, i∈[1,N] and ABi∈R. C×S×S By changing the preset size, the number of anchor boxes can be altered, allowing for changes in the focus on features at different scales within the feature map, especially small details that affect image classification results. Each first feature sub-image is input into a feature attention layer for processing, yielding a processed second feature sub-image. These second feature sub-images are then input into a feature stitching layer for concatenation, resulting in a processed second fundus feature map. This allows for attention processing of feature maps at different finer granularities, further improving the accuracy of image classification results.

[0080] Based on the above technical solution, "inputting each first feature sub-image into a feature attention layer for feature attention processing to obtain a processed second feature sub-image" may include: for each first feature sub-image, inputting the first feature sub-image into a second average pooling layer for average pooling processing to obtain an average pooled third feature sub-image; inputting the first feature sub-image into a second max pooling layer for max pooling processing to obtain a max pooled fourth feature sub-image; inputting the third and fourth feature sub-images into a feature convolution layer for superposition and convolution processing to obtain a processed fifth feature sub-image; and inputting the fifth feature sub-image and the first feature sub-image into a second residual processing layer for residual processing to obtain a processed second feature sub-image.

[0081] The feature attention layer may include a second average pooling layer, a second max pooling layer, a feature convolutional layer, and a second residual processing layer. Figure 9 A schematic diagram of a feature attention layer is provided. The third feature sub-image can be obtained by average pooling the first feature sub-image. The fourth feature sub-image can be obtained by max pooling the first feature sub-image. The fifth feature sub-image can be the feature sub-image obtained by superimposing the two feature sub-images.

[0082] Specifically, see Figure 9 For each first feature sub-map (ABi), the first feature sub-map (ABi) can be input into the second average pooling layer along the channel direction for average pooling processing to obtain the average pooled third feature sub-map (F). avg (ABi)); The first feature sub-map (ABi) can be input into the second max pooling layer along the channel direction for max pooling processing to obtain the fourth feature sub-map (F) after max pooling. max (ABi)); the third feature subgraph (F) avg (ABi)) and the fourth feature subgraph (F max (ABi)) is input to the feature convolutional layer and stacked and convolved along the channel direction to obtain the processed fifth feature sub-map (Conv([F)). avg (ABi); F max (ABi)])), weight the fifth feature subgraph, and then calculate the weighted result (σ(Conv([F avg (ABi); F max (ABi)]))) and the first feature sub-map (ABi) are input into the second residual processing layer for element-wise multiplication to obtain the processed second feature sub-map (ABi').

[0083] The following are embodiments of the fundus image classification device provided in this invention. This device and the fundus image classification methods in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the fundus image classification device, please refer to the embodiments of the fundus image classification methods described above.

[0084] Example 3

[0085] Figure 10 This is a schematic diagram of the structure of a fundus image classification device provided in Embodiment 3 of the present invention. Figure 10 As shown, the device includes: a first image acquisition module 310, a second image determination module 320, an image classification module 330, and a classification result determination module 340.

[0086] The system includes a first image acquisition module 310 for acquiring a first fundus image to be classified; a second image determination module 320 for preprocessing the first fundus image to obtain a preprocessed second fundus image; an image classification module 330 for inputting the second fundus image into a pre-trained image classification network model for image classification, wherein the image classification network model performs fine-grained feature attention processing on the second fundus image and classifies it based on the processing results; and a classification result determination module 340 for determining the classification result corresponding to the first fundus image based on the output of the image classification network model.

[0087] The technical solution of this invention preprocesses the first fundus image to be classified, and then inputs the preprocessed second fundus image into a pre-trained image classification network model for image classification. The image classification network model performs fine-grained feature attention processing on the input second fundus image and classifies it based on the processing results. This can improve the attention of useful information and suppress useless information, so that the image classification network model can automatically output more accurate classification results, improve the accuracy of image classification, and achieve automatic image classification without human intervention, thereby improving the efficiency of image classification.

[0088] Optionally, the second image determination module 320 is specifically used to: perform Gaussian filtering on the first fundus image to obtain a filtered third fundus image; and determine the second fundus image based on the first fundus image, the third fundus image, and a preset bias pixel value.

[0089] Optionally, the image classification network model includes: a feature extraction sub-model, an attention processing sub-model, and a classification processing sub-model;

[0090] Image classification module 330 may include:

[0091] The feature extraction submodule is used to input the second fundus image into the feature extraction submodel for feature extraction, and obtain the extracted first fundus feature map;

[0092] The attention processing submodule is used to input the first fundus feature map into the attention processing submodel for fine-grained attention processing to obtain the processed second fundus feature map;

[0093] The classification result determination submodule is used to input the second fundus feature map into the classification processing submodel for fully connected processing to determine the classification result corresponding to the first fundus image.

[0094] Optionally, the attention processing sub-model includes: a feature mapping unit and a fine-grained attention unit;

[0095] The attention processing submodule may include:

[0096] The feature mapping unit is used to input the first fundus feature map into the feature mapping unit for feature mapping to obtain the third fundus feature map after feature mapping.

[0097] The attention processing unit is used to input the third fundus feature map into the fine-grained attention unit for fine-grained attention processing to obtain the processed second fundus feature map.

[0098] Optionally, the feature mapping unit includes: a first average pooling layer, a first max pooling layer, a first multilayer perceptron, a second multilayer perceptron, an activation processing layer, and a first residual processing layer;

[0099] The feature mapping unit is specifically used for: inputting the first fundus feature map into the first average pooling layer for average pooling processing to obtain the average pooled fourth fundus feature map; inputting the fourth fundus feature map into the first multilayer perceptron for processing to obtain the processed fifth fundus feature map; inputting the first fundus feature map into the first max pooling layer for max pooling processing to obtain the max pooled sixth fundus feature map; inputting the sixth fundus feature map into the second multilayer perceptron for processing to obtain the processed seventh fundus feature map; inputting the fifth and seventh fundus feature maps into the activation processing layer for activation processing to obtain the activated eighth fundus feature map; and inputting the first and eighth fundus feature maps into the first residual processing layer for residual processing to obtain the feature-mapped third fundus feature map.

[0100] Optionally, the fine-grained attention unit includes: a feature segmentation layer, a feature attention layer, and a feature concatenation layer;

[0101] Attention processing unit may include:

[0102] The feature segmentation subunit is used to input the third fundus feature map into the feature segmentation layer for feature map segmentation processing to obtain multiple first feature sub-maps in the third fundus feature map;

[0103] The attention processing subunit is used to input each first feature sub-map into the feature attention layer for feature attention processing to obtain the processed second feature sub-map.

[0104] The feature stitching sub-unit is used to input each second feature sub-image into the feature stitching layer for feature sub-image stitching processing to obtain the processed second fundus feature map.

[0105] Optionally, the feature attention layer includes: a second average pooling layer, a second max pooling layer, a feature convolutional layer, and a second residual processing layer;

[0106] The attention processing subunit is specifically used for: for each first feature sub-image, inputting the first feature sub-image into the second average pooling layer for average pooling processing to obtain the average pooled third feature sub-image; inputting the first feature sub-image into the second max pooling layer for max pooling processing to obtain the max pooled fourth feature sub-image; inputting the third and fourth feature sub-images into the feature convolution layer for stacking and convolution processing to obtain the processed fifth feature sub-image; and inputting the fifth feature sub-image and the first feature sub-image into the second residual processing layer for residual processing to obtain the processed second feature sub-image.

[0107] The fundus image classification device provided in the embodiments of the present invention can execute the fundus image classification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the fundus image classification method.

[0108] It is worth noting that in the embodiments of the fundus image classification device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of the present invention.

[0109] Example 4

[0110] Figure 11 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0111] like Figure 11As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0112] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0113] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as fundus image classification methods.

[0114] In some embodiments, the fundus image classification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the fundus image classification method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the fundus image classification method by any other suitable means (e.g., by means of firmware).

[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0120] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A fundus image classification method characterized by, The method comprises the following steps: obtaining a first fundus image to be classified; preprocessing the first fundus image to obtain a second fundus image after preprocessing; inputting the second fundus image into a pre-trained image classification network model for image classification, wherein the image classification network model is used for attention processing of fine-grained features of the second fundus image, and classification is performed based on the processing result; determining the classification result corresponding to the first fundus image according to the output of the image classification network model; wherein the image classification network model comprises a feature extraction sub-model and an attention processing sub-model; the attention processing sub-model comprises a feature mapping unit and a fine-grained attention unit; the fine-grained attention unit comprises a feature segmentation layer, a feature attention layer and a feature splicing layer; wherein the first fundus feature map is obtained by feature extraction of the second fundus image by the feature extraction sub-model; the third fundus feature map is obtained by feature mapping of the first fundus feature map by the feature mapping unit; wherein the plurality of first feature sub-maps in the third fundus feature map are obtained by feature map segmentation processing of the third fundus feature map by the feature segmentation layer; the second feature sub-map is obtained by feature attention processing of each first feature sub-map by the feature attention layer; and the second fundus feature map is obtained by feature sub-map splicing processing of each second feature sub-map by the feature splicing layer.

2. The method of claim 1, wherein, The preprocessing of the first fundus image to obtain the second fundus image after preprocessing comprises: performing Gaussian filtering on the first fundus image to obtain a third fundus image after filtering; determining the second fundus image based on the first fundus image, the third fundus image and a preset bias pixel value.

3. The method of claim 1, wherein, The image classification network model comprises a feature extraction sub-model, an attention processing sub-model and a classification processing sub-model; The inputting of the second fundus image into the pre-trained image classification network model for image classification comprises: inputting the second fundus image into the feature extraction sub-model for feature extraction to obtain the extracted first fundus feature map; inputting the first fundus feature map into the attention processing sub-model for fine-grained attention processing to obtain the processed second fundus feature map; inputting the second fundus feature map into the classification processing sub-model for full connection processing to determine the classification result corresponding to the first fundus image.

4. The method of claim 3, wherein, The attention processing sub-model comprises a feature mapping unit and a fine-grained attention unit; The inputting of the first fundus feature map into the attention processing sub-model for fine-grained attention processing to obtain the processed second fundus feature map comprises: inputting the first fundus feature map into the feature mapping unit for feature mapping to obtain the third fundus feature map after feature mapping; inputting the third fundus feature map into the fine-grained attention unit for fine-grained attention processing to obtain the processed second fundus feature map.

5. The method of claim 4, wherein, The feature mapping unit comprises a first average pooling layer, a first maximum pooling layer, a first multilayer perceptron, a second multilayer perceptron, an activation processing layer and a first residual processing layer; The first fundus feature map is input into the feature mapping unit for feature mapping to obtain a third fundus feature map after feature mapping. The first fundus feature map is input into the first average pooling layer for average pooling processing to obtain a fourth fundus feature map after average pooling. The fourth fundus feature map is input into the first multilayer perceptron for processing to obtain a fifth fundus feature map after processing. The first fundus feature map is input into the first maximum pooling layer for maximum pooling processing to obtain a sixth fundus feature map after maximum pooling. The sixth fundus feature map is input into the second multilayer perceptron for processing to obtain a seventh fundus feature map after processing. The fifth fundus feature map and the seventh fundus feature map are input into the activation processing layer for activation processing to obtain an eighth fundus feature map after activation. The first fundus feature map and the eighth fundus feature map are input into the first residual processing layer for residual processing to obtain a third fundus feature map after feature mapping.

6. The method of claim 4, wherein, The fine-grained attention unit comprises a feature segmentation layer, a feature attention layer and a feature splicing layer; The third fundus feature map is input into the fine-grained attention unit for fine-grained attention processing to obtain a second fundus feature map after processing. The third fundus feature map is input into the feature segmentation layer for feature map segmentation processing to obtain a plurality of first feature sub-maps in the third fundus feature map; Each first feature sub-map is input into the feature attention layer for feature attention processing to obtain a second feature sub-map after processing. Each second feature sub-map is input into the feature splicing layer for feature sub-map splicing processing to obtain a second fundus feature map after processing.

7. The method of claim 6, wherein, The feature attention layer comprises a second average pooling layer, a second maximum pooling layer, a feature convolution layer and a second residual processing layer; Each first feature sub-map is input into the feature attention layer for feature attention processing to obtain a second feature sub-map after processing. For each first feature sub-map, the first feature sub-map is input into the second average pooling layer for average pooling processing to obtain a third feature sub-map after average pooling. The first feature sub-map is input into the second maximum pooling layer for maximum pooling processing to obtain a fourth feature sub-map after maximum pooling. The third feature sub-map and the fourth feature sub-map are input into the feature convolution layer for superposition and convolution processing to obtain a fifth feature sub-map after processing. The fifth feature sub-map and the first feature sub-map are input into the second residual processing layer for residual processing to obtain a second feature sub-map after processing.

8. An ocular fundus image classification apparatus characterized by comprising: It comprises: A first image acquisition module is configured to acquire a first fundus image to be classified; A second image determination module is configured to preprocess the first fundus image to obtain a second fundus image after preprocessing. The image classification module is configured to input the second fundus image into a pre-trained image classification network model to perform image classification, wherein the image classification network model is configured to perform attention processing on fine-grained features of the second fundus image and perform classification based on a processing result. The classification result determination module is configured to determine a classification result corresponding to the first fundus image based on an output of the image classification network model. The image classification network model includes a feature extraction sub-model and an attention processing sub-model, the attention processing sub-model includes a feature mapping unit and a fine-grained attention unit, and the fine-grained attention unit includes a feature segmentation layer, a feature attention layer, and a feature concatenation layer. The first fundus feature map is obtained by performing feature extraction on the second fundus image by using the feature extraction sub-model, and the third fundus feature map is obtained by performing feature mapping on the first fundus feature map by using the feature mapping unit. The first feature sub-maps in the third fundus feature map are obtained by performing feature map segmentation processing on the third fundus feature map by using the feature segmentation layer, the second feature sub-maps are obtained by performing feature attention processing on each first feature sub-map by using the feature attention layer, and the second fundus feature map is obtained by performing feature sub-map concatenation processing on each second feature sub-map by using the feature concatenation layer.

9. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the fundus image classification method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to execute the fundus image classification method of any one of claims 1-7 when executed.

Citation Information

Patent Citations

  • Eye fundus image retinal vessel segmentation method based on mixed attention mechanism

    CN112132817A

  • Product defect detection method, apparatus and system

    WO2021135331A1