Eye fundus image quality evaluation method, system and device and storage medium

By adopting a dual-current cross-fusion image classification network in fundus image quality evaluation, the problem of difficulty in evaluating intermediate-state quality images is solved, the accurate identification and utilization of intermediate-state quality images is achieved, and the accuracy and efficiency of fundus image quality evaluation is improved.

CN119941736AActive Publication Date: 2025-05-06NINGBO INST OF MATERIALS TECH & ENG CHINESE ACAD OF SCI +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510431365.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing fundus image quality assessment methods are difficult to accurately evaluate intermediate-state quality images, and are prone to misclassification as unavailable images, thereby wasting patient data resources.

Method used

The image quality classification network based on dual-stream cross-fusion is adopted. Through the dual-stream network module, cross-fusion module, residual channel attention module and classification module, the correlation characteristics of high-quality and low-quality images are extracted, and the cross-fusion and attention mechanisms are used to accurately identify intermediate-state quality images.

Benefits of technology

The classification accuracy of intermediate-state quality images is improved, and the images are avoided misclassification as unavailable images, thereby effectively utilizing patient data resources and improving the quality evaluation effect of fundus images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941736A_ABST
    Figure CN119941736A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of image processing, and provides a fundus image quality evaluation method, system and device and a storage medium, and a quality evaluation model of the fundus image quality evaluation system comprises a double-flow network module, a cross fusion module, a first RCAM module, a second RCAM module and a classification module. The double-flow network module comprises a first network and a second network; the high-quality image correlation features extracted by the first network are used for subsequently judging whether the fundus image is a high-quality image, the low-quality image correlation features extracted by the second network are used for subsequently judging whether the fundus image is a low-quality image, and the arrangement of the double-flow network module enables the method to accurately recognize an intermediate-state quality image. And waste caused by abandoning of the type of images is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and relates to a fundus image quality assessment method, system, device and storage medium. Background Art

[0002] Fundus image quality assessment is a vital part of medical imaging, especially in ophthalmic diagnosis, where accurate image quality plays a vital role in disease screening and the formulation of treatment plans. Fundus images are widely used to diagnose a variety of ophthalmic diseases, such as retinal diseases, diabetic retinopathy, macular degeneration, and glaucoma. Due to differences in equipment, shooting conditions, imaging technology, and doctor's operating level, the quality of fundus images varies greatly. Some low-quality fundus images may affect the identification and accurate diagnosis of lesions due to blur, noise, or viewing angle problems, so these images need to be identified and excluded. Another part of the fundus images with lower quality still contain valuable pathological information, and their quality can be improved through reasonable image enhancement processing, which can then be used for effective diagnosis.

[0003] Existing methods for fundus image quality assessment mainly rely on manual inspection or traditional image quality scoring systems. These methods usually rely on the doctor's experience, are inefficient, and are easily affected by human subjective factors. With the continuous development of deep learning technology, automated image quality assessment methods based on convolutional neural networks (CNN) have gradually been applied. These methods can automatically learn image features to score quality and automatically detect low-quality images, significantly improving the accuracy and efficiency of image quality assessment. However, the existing image quality assessment methods are not effective for fundus images that are affected by degradation features but to a lesser extent, contain pathological structures that can be distinguished by the naked eye, and can be used for clinical diagnosis through subsequent image enhancement (i.e., intermediate quality images). These images are easily classified as unusable images, thereby wasting part of the patient's data resources. Summary of the invention

[0004] Aiming at the problem that it is difficult to define intermediate quality images in existing image quality assessment methods, the present invention proposes a fundus image quality assessment method, system, device and storage medium.

[0005] The present invention is achieved through the following technical solutions: A fundus image quality assessment system, comprising: A quality assessment model is used to classify the quality of fundus images to obtain quality categories; The quality assessment model includes: A dual-stream network module, comprising a first network and a second network for respectively extracting high-quality image-related features and low-quality image-related features in fundus images; the first network and the second network both comprise a plurality of densely connected blocks connected in sequence; the outputs of the densely connected blocks of the first network are concatenated as the first output of the first network, and the output of the last densely connected block is the second output of the first network; the outputs of the densely connected blocks of the second network are concatenated as the first output of the second network, and the output of the last densely connected block is the second output of the second network; A cross-fusing module (CFM), configured to perform cross-fusing processing on the first output of the first network and the first output of the second network to obtain a first CFM output and a second CFM output; The first RCAM module (Residual Channel Attention Module) is used to focus on the first CFM output and the second output of the first network to obtain the first classification feature; A second RCAM module; used for focusing attention on the second CFM output and the second network second output to obtain a second classification feature; The classification module is used to obtain a quality category according to the first classification feature and the second classification feature.

[0006] Preferably, the fundus image quality assessment system also includes a Gabor filtering module; the Gabor filtering module is used to perform Gabor filtering on the fundus image to obtain a Gabor filter image; the input image of the first network is the fundus image, and the input image of the second network is the Gabor filter image.

[0007] Preferably, the cross-fusion module includes: A first channel is used to extract local features of a first output of the first network; A first cross-fusion summing module, used for adding the output of the first channel to the first output of the second network; The first SE (channel attention, Squeeze-and-Excitation) module (SEM) is used to perform feature selection on the output of the first cross-fusion summation module; A first post-processing module, used for adjusting the number of characteristic channels of the output of the first SE module so that the number of characteristic channels of the output of the first SE module is consistent with the number of characteristic channels of the first output of the first network, and obtaining a first CFM output; A second channel is used to extract local features of the first output of the second network; A second cross-fusion summing module is used to perform addition processing on the output of the second channel and the first output of the first network; The second SE module is used to perform feature selection on the output of the second cross-fusion summation module; The second post-processing module is used to adjust the number of feature channels of the output of the second SE module so that the number of feature channels of the output of the second SE module is consistent with the number of feature channels of the first output of the second network, thereby obtaining a second CFM output.

[0008] Furthermore, the first channel includes: a first cross-fusion convolution layer, a first cross-fusion batch normalization (Batch Normalization, BN) layer and a first cross-fusion ReLU (Rectified Linear Unit) activation function module; the second channel includes: a second cross-fusion convolution layer, a second cross-fusion batch normalization layer and a second cross-fusion ReLU activation function module; The first output of the first network is processed in sequence by a first cross-fusion convolutional layer, a first cross-fusion batch normalization layer, and a first cross-fusion ReLU activation function module to obtain an output of a first channel; The first output of the second network is processed in sequence by a second cross-fusion convolutional layer, a second cross-fusion batch normalization layer, and a second cross-fusion ReLU activation function module to obtain an output of a second channel.

[0009] Preferably, the first RCAM module comprises: A first residual convolution layer, used to perform a convolution operation on the first CFM output; A first residual summing module, used for performing feature element-level addition of the output of the first residual convolutional layer and the second output of the first network; A second residual convolution layer, used to perform a convolution operation on the output of the first residual summation module; A first Sigmoid activation function module is used to process the output of the second residual convolution layer using the Sigmoid activation function to obtain a first attention weight; The third residual convolution layer is used to perform a convolution operation on the first CFM output; A first multiplication module, used for multiplying the first attention weight and the output of the third residual convolution layer to obtain a first selected feature; The first global pooling (GP) layer is used to perform pooling on the first CFM output; A fourth residual convolution layer, used to perform a convolution operation on the output of the first global pooling layer; A first residual ReLU activation function module, used to process the output of the fourth residual convolution layer using a ReLU activation function to obtain a first channel attention score; The second multiplication module is used to multiply the first selected feature by the first channel attention score to obtain a first classification feature.

[0010] Further, the second RCAM module includes: A fifth residual convolution layer, used to perform a convolution operation on the second CFM output; A second residual summing module, used for performing feature element-level addition of the output of the fifth residual convolutional layer and the second output of the second network; A sixth residual convolution layer, used for performing a convolution operation on the output of the second residual summation module; A second Sigmoid activation function module is used to process the output of the sixth residual convolution layer using a Sigmoid activation function to obtain a second attention weight; A seventh residual convolution layer, used for performing a convolution operation on the second CFM output; A third multiplication module is used to multiply the second attention weight and the output of the seventh residual convolution layer to obtain a second selected feature; A second global pooling layer, used for performing pooling processing on the second CFM output; An eighth residual convolution layer, used to perform a convolution operation on the output of the second global pooling layer; A second residual ReLU activation function module, used to process the output of the eighth residual convolution layer using the ReLU activation function to obtain a second channel attention score; The fourth multiplication module is used to multiply the second selected feature by the second channel attention score to obtain a second classification feature.

[0011] Preferably, the fundus image quality assessment system further includes a data augmentation module; The data augmentation module is used to expand the fundus images of a few quality categories in the training data set before training the quality assessment model.

[0012] A fundus image quality assessment method uses the fundus image quality assessment system as described above to classify the quality of the fundus image to obtain a quality category.

[0013] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the fundus image quality assessment method as described above when executing the computer program.

[0014] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the fundus image quality assessment method as described above is implemented.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The fundus image quality assessment system of the present invention is an image quality classification network based on dual-stream cross-fusion, and its quality assessment model includes a dual-stream network module, a cross-fusion module, a residual channel attention module (i.e., a first RCAM module and a second RCAM module) and a classification module to achieve accurate identification of image quality categories. The first network and the second network in the dual-stream network module are respectively used to extract high-quality image-related features and low-quality image-related features in the fundus image. The high-quality image-related features extracted by the first network are used to subsequently determine whether the fundus image is a high-quality image, and the low-quality image-related features extracted by the second network are used to subsequently determine whether the fundus image is a low-quality image. When the former judgment result is a non-high-quality image and the latter judgment result is a non-low-quality image, the fundus image is considered to be an intermediate quality image. Therefore, the setting of the dual-stream network module enables the present invention to accurately identify intermediate quality images and avoid waste caused by abandonment of images of this category. The cross-fusion module performs cross-fusion between the first network and the second network to extract useful information from the fundus image to the greatest extent, which can solve the problem of no information exchange and no feature interaction between the first network and the second network, and realize efficient fusion of information between channels and feature enhancement; the features output by the cross-fusion module couple the useful information of the fundus image itself, aiming to improve the robustness of the quality assessment model classification. The residual channel attention module enhances the feature selection capability with the help of the attention mechanism. The module can complete feature selection, generate the final classification features, and improve the stability of the quality assessment model. The system of the present invention can be used for fundus image quality analysis and intelligent medical image assessment of premature infants.

[0016] Furthermore, most existing classification models do not pay attention to the structural features of fundus images, lack the ability to characterize structural features, and need to assume in advance which degradation form the fundus image is affected by, resulting in degradation features. The sensitivity to fundus images containing multiple degradation features is insufficient. For this problem, the first network of the present invention uses the original fundus image as the input image, and the second network uses the Gabor filter image of the original fundus image as the input image. The Gabor filter image can highlight the structural features of the fundus image. After inputting the second network, it will be easier and more sufficient to extract the structural features, so that the fundus image can be automatically judged to what extent it is affected by multiple degradation forms, and the number of mixed degradation features in the fundus image can be judged. The sensitivity to fundus images containing multiple degradation features is higher. Therefore, the present invention can better extract the structural features of the fundus image, identify multiple degradation features, and thus improve the accuracy of fundus image quality classification.

[0017] Furthermore, the present invention provides a data augmentation module for expanding fundus images of a few quality categories, thereby generating diverse training data, significantly improving the generalization ability of the quality assessment model, and enabling it to maintain high performance in various practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 A framework for the quality assessment model of the present invention; Figure 2 It is the framework of the cross-fusion module of the present invention; Figure 3 It is the framework of the residual channel attention module of the present invention. DETAILED DESCRIPTION

[0020] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention.

[0021] It should be noted that the process equipment or devices not specifically specified in the following embodiments are all conventional equipment or devices in the art.

[0022] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. Moreover, unless otherwise specified, the numbering of each method step is only a convenient tool for identifying each method step, and is not intended to limit the order of arrangement of each method step or to define the scope of the present invention. Changes or adjustments in their relative relationships should also be regarded as the scope of the present invention without substantially changing the technical content.

[0023] The fundus image quality assessment system of the present invention comprises: A quality assessment model is used to classify the quality of the input fundus image to obtain a quality category; Among them, Figure 1 As shown, the quality assessment model includes: A dual-stream network module, comprising a first network and a second network, respectively used to extract high-quality image-related features and low-quality image-related features in fundus images; the first network and the second network respectively comprise a plurality of densely connected blocks connected in sequence; the outputs of the densely connected blocks in the first network are concatenated in a feature channel dimension as the first output of the first network, and the output of the last densely connected block is the second output of the first network; the outputs of the densely connected blocks in the second network are concatenated in a feature channel dimension as the first output of the second network, and the output of the last densely connected block is the second output of the second network; A cross-fusion module, used for performing cross-fusion processing on the first output of the first network and the first output of the second network to obtain a first CFM output and a second CFM output corresponding to the first network and the second network respectively; The residual channel attention module includes a first RCAM module and a second RCAM module; the first RCAM module takes the first CFM output and the second output of the first network as input, performs attention focusing processing, and obtains a first classification feature; the second RCAM module takes the second CFM output and the second network second output as input, performs attention focusing processing, and obtains a second classification feature; The classification module is used to obtain a quality category according to the first classification feature and the second classification feature.

[0024] In the present invention, the first network and the second network are both densely connected convolutional networks (DenseNet), and the high-quality image-related features and the low-quality image-related features both include degradation features and structural features.

[0025] In a preferred embodiment of the present invention, the fundus image quality assessment system further includes an image acquisition module and a result output module; The image acquisition module is used to acquire a fundus image and input the fundus image into a quality assessment model; The result output module is used to output the quality category.

[0026] In a preferred embodiment of the present invention, the fundus image quality assessment system also includes a Gabor filtering module; the Gabor filtering module is used to perform Gabor filtering on the fundus image to obtain a Gabor filter image; the input image of the first network is the fundus image, and the input image of the second network is the Gabor filter image.

[0027] The first network of the present invention uses the original fundus image as the input image, and the second network uses the Gabor filter image of the original fundus image as the input image. The Gabor filter image can highlight the structural features of the fundus image. After inputting into the second network, it will be easier and more complete to extract the structural features.

[0028] like Figure 2 As shown, the cross-fusion module of the present invention includes: A first channel is used to extract local features of a first output of the first network; A first cross-fusion summing module, used for adding the output of the first channel to the first output of the second network; A first SE module, used for performing feature selection on the output of the first cross-fusion summation module; A first post-processing module, used for adjusting the number of characteristic channels of the output of the first SE module so that the number of characteristic channels of the output of the first SE module is consistent with the number of characteristic channels of the first output of the first network, and obtaining a first CFM output; A second channel is used to extract local features of the first output of the second network; A second cross-fusion summing module is used to perform addition processing on the output of the second channel and the first output of the first network; The second SE module is used to perform feature selection on the output of the second cross-fusion summation module; The second post-processing module is used to adjust the number of feature channels of the output of the second SE module so that the number of feature channels of the output of the second SE module is consistent with the number of feature channels of the first output of the second network, thereby obtaining a second CFM output.

[0029] Wherein, the first channel includes: a first cross-fusion convolution layer, a first cross-fusion batch normalization layer and a first cross-fusion ReLU activation function module; the second channel includes: a second cross-fusion convolution layer, a second cross-fusion batch normalization layer and a second cross-fusion ReLU activation function module; The first output of the first network is processed in sequence by a first cross-fusion convolutional layer, a first cross-fusion batch normalization layer, and a first cross-fusion ReLU activation function module to obtain an output of a first channel; The first output of the second network is processed in sequence by a second cross-fusion convolutional layer, a second cross-fusion batch normalization layer, and a second cross-fusion ReLU activation function module to obtain an output of a second channel.

[0030] Among them, the first post-processing module includes: a third cross-fusion convolution layer, a third cross-fusion batch normalization layer and a third cross-fusion ReLU activation function module; the second post-processing module includes: a fourth cross-fusion convolution layer, a fourth cross-fusion batch normalization layer and a fourth cross-fusion ReLU activation function module; The output of the first SE module is processed in sequence by a third cross-fusion convolutional layer, a third cross-fusion batch normalization layer, and a third cross-fusion ReLU activation function module to obtain a first CFM output; The output of the second SE module is processed in sequence by a fourth cross-fusion convolutional layer, a fourth cross-fusion batch normalization layer and a fourth cross-fusion ReLU activation function module to obtain a second CFM output.

[0031] like Figure 3 , which is a schematic diagram of one of the RCAM modules, is applicable to the first RCAM module and the second RCAM module.

[0032] The first RCAM module of the present invention comprises: A first residual convolution layer, used to perform a convolution operation on the first CFM output; A first residual summing module, used for performing feature element-level addition of the output of the first residual convolutional layer and the second output of the first network; A second residual convolution layer, used to perform a convolution operation on the output of the first residual summation module; A first Sigmoid activation function module is used to process the output of the second residual convolution layer using the Sigmoid activation function to obtain a first attention weight; The third residual convolution layer is used to perform a convolution operation on the first CFM output; A first multiplication module, used for multiplying the first attention weight and the output of the third residual convolution layer to obtain a first selected feature; A first global pooling layer, used for performing pooling processing on the first CFM output; A fourth residual convolution layer, used to perform a convolution operation on the output of the first global pooling layer; A first residual ReLU activation function module, used to process the output of the fourth residual convolution layer using a ReLU activation function to obtain a first channel attention score; The second multiplication module is used to multiply the first selected feature by the first channel attention score to obtain a first classification feature.

[0033] The second RCAM module comprises: A fifth residual convolution layer, used to perform a convolution operation on the second CFM output; A second residual summing module, used for performing feature element-level addition of the output of the fifth residual convolutional layer and the second output of the second network; A sixth residual convolution layer, used for performing a convolution operation on the output of the second residual summation module; A second Sigmoid activation function module is used to process the output of the sixth residual convolution layer using a Sigmoid activation function to obtain a second attention weight; A seventh residual convolution layer, used for performing a convolution operation on the second CFM output; A third multiplication module is used to multiply the second attention weight and the output of the seventh residual convolution layer to obtain a second selected feature; A second global pooling layer, used for performing pooling processing on the second CFM output; An eighth residual convolution layer, used to perform a convolution operation on the output of the second global pooling layer; A second residual ReLU activation function module, used to process the output of the eighth residual convolution layer using the ReLU activation function to obtain a second channel attention score; The fourth multiplication module is used to multiply the second selected feature by the second channel attention score to obtain a second classification feature.

[0034] The output of the cross-fusion module of the present invention generates the final classification features in the residual channel attention module through residual connection to improve the stability of the quality assessment model. The residual channel attention module effectively avoids the gradient vanishing problem in the quality assessment model training process through residual connection, ensuring the effective training of the quality assessment model; with the help of the attention mechanism, the feature selection ability is enhanced, and the module can complete feature selection and generate the final classification features.

[0035] The classification module described in the present invention maps the first classification feature and the second classification feature to the Bernoulli distribution of different quality categories respectively to obtain the Bernoulli probability, and obtains the quality category according to the Bernoulli probability.

[0036] like Figure 1 As shown, the classification module of the present invention specifically includes: A third global pooling layer is used to perform global pooling processing on the first classification feature; A fourth global pooling layer, used for performing global pooling processing on the second classification feature; The first fully connected layer (FC) maps the first classification feature of the global pooling to the high-quality Bernoulli distribution classifier to obtain the first Bernoulli probability; The second fully connected layer is used to map the globally pooled second classification features to the low-quality Bernoulli distribution classifier to obtain the second Bernoulli probability; The quality category judgment module is used to obtain the quality category according to the first Bernoulli probability and the second Bernoulli probability.

[0037] In order to better train the quality assessment model, increase the number of fundus images in minority quality categories, and balance the data set, the present invention can also set a data augmentation module in the fundus image quality assessment system. The data augmentation module is used to expand the fundus images in minority quality categories in the training data set before training the quality assessment model. The minority quality category is relative to the majority quality category. When the ratio of the number of fundus images in one quality category to the number of fundus images in another quality category is less than 0.5, the quality category is considered to be a minority quality category. By expanding the fundus images in minority quality categories through the data augmentation module, the problem of heavy workload for doctors caused by relying on doctor expansion can be avoided.

[0038] In order to perform visual analysis on degradation features and structural features, the fundus image quality assessment system of the present invention also includes a gradient-weighted class activation mapping (GradCAM) visualization processing module; the GradCAM visualization processing module is used to generate a heat map by calculating the gradient of the first classification feature relative to the output of the third residual convolution layer and the gradient of the second classification feature relative to the output of the seventh residual convolution layer.

[0039] By combining GradCAM visualization analysis, the present invention provides an intuitive feature distribution explanation, which not only displays the degradation features and structural features that the quality assessment model focuses on, but also helps doctors understand the classification basis of the quality assessment model and improves the credibility of the classification results.

[0040] The fundus image quality assessment method of the present invention comprises: Input the fundus image into a pre-trained quality assessment model for quality classification to obtain a quality category; The quality assessment model is the quality assessment model in the fundus image quality assessment system as described above.

[0041] Example 1 The fundus image quality assessment system of this embodiment includes: an image acquisition module, a quality assessment model and a result output module; The image acquisition module is used to acquire the fundus image and input the acquired fundus image into the quality assessment model; The quality assessment model is used to perform quality classification on the input fundus image to obtain a quality category; The result output module is used to output the quality category.

[0042] The present invention uses a data set with quality categories annotated by doctors to train a quality assessment model. The trained quality assessment model takes fundus images as input to achieve accurate assessment of fundus image quality.

[0043] The dataset of doctor-labeled quality categories is obtained by the following method: First, the original fundus images are collected and preliminarily cleaned and preprocessed, including archiving fundus images with the same number, unifying the naming format of fundus images, and classifying the left and right eye image data to ensure the unity and integrity of the fundus images; the number is the shooting date of the fundus image; Then, the fundus images are annotated and the quality category of each fundus image is confirmed to form an initial data set.

[0044] In this embodiment, the fundus image quality assessment system further includes a data augmentation module, which is used to expand fundus images of a few quality categories in a training data set before training the quality assessment model.

[0045] Before training the quality assessment model, the present invention designs a data augmentation module to process the initial data set to improve the performance of the quality assessment model. The core of the data augmentation module is to use a data augmentation algorithm to expand the fundus images of a few quality categories in the initial data set to balance the number of fundus images of each quality category in the initial data set.

[0046] In the present invention, the specific steps of the data augmentation module to expand the fundus images of a few quality categories are as follows: 1) Analyze the number of fundus images of each quality category in the initial data set and identify the minority quality category and the majority quality category. For example, if the ratio of the number of fundus images in the low-quality category to the number of fundus images in the high-quality category is less than 0.5, the low-quality category is considered to be the minority quality category and the fundus images in the low-quality category need to be expanded. For the identified minority quality category, apply the image data augmentation method to generate more fundus images. Commonly used image data augmentation methods include but are not limited to rotation, scaling, translation, flipping, color transformation and other operations to generate new image data.

[0047] 2) The augmented fundus images of the minority quality category are merged with the initial dataset to form a new, balanced dataset. Doctors will recheck the new dataset to determine whether the augmented fundus images match the minority quality category to ensure data quality.

[0048] Through the above steps, the present invention achieves the technical effect of effectively increasing the number of fundus images in a minority quality category and balancing the data set.

[0049] The quality assessment model described in this embodiment is an image quality classification network based on dual-stream cross-fusion feature extraction, specifically including a dual-stream network module, a cross-fusion module, a residual channel attention module, a classification module and a visualization processing module, which is used to extract diverse degradation features and structural features to achieve accurate identification of image quality categories.

[0050] The present invention designs a dual-stream network module, and proposes a new quality classification strategy based on the dual-stream network module to assist in the accurate classification of intermediate-quality fundus images; designs a cross-fusion module to fuse high-quality image-related features and low-quality image-related features; designs a residual channel attention module, uses residual connections to avoid gradient vanishing in the quality assessment model training process, completes feature selection with the help of the attention mechanism, and generates the final classification features. The present invention uses the fully connected layer to output Bernoulli distribution to predict the quality category, and uses GradCAM to visualize the degradation features and structural features, providing an interpretable basis for the feature distribution for the fundus image quality classification results.

[0051] The following is a detailed introduction to each module of the quality assessment model.

[0052] (1) Two-stream network module to assist in intermediate state quality classification The present invention introduces a dual-stream network module for assisting intermediate quality classification, and the dual-stream network module includes two networks, namely, a first network (high-quality classifier) ​​and a second network (low-quality classifier), i.e., a dual-stream network. The first network and the second network are used to extract high-quality image-related features and low-quality image-related features in fundus images, respectively, and a DenseNet network is used as the backbone network of the first network and the second network. The high-quality image-related features and the low-quality image-related features both include degradation features and structural features. The degradation features include blur and noise, and the structural features include vascular structures and fundus pathological structures.

[0053] In this embodiment, the fundus image quality assessment system further includes a Gabor filtering module, which is used to perform Gabor filtering on the fundus image acquired by the image acquisition module to obtain a Gabor filter image; the input image of the first network is the fundus image, and the input image of the second network is the Gabor filter image. The Gabor filter image can highlight the structural features of the fundus image, and it will be easier and more sufficient to extract the structural features after inputting it into the second network.

[0054] The first network and the second network respectively include a plurality of densely connected blocks connected in sequence; the features of the outputs of each densely connected block in the first network are concatenated in the feature channel dimension as the first output of the first network, and the output of the last densely connected block is used as the second output of the first network; the features of the outputs of each densely connected block in the second network are concatenated in the feature channel dimension as the first output of the second network, and the output of the last densely connected block is used as the second output of the second network.

[0055] (2) Cross-fusion module The present invention provides a cross-fusion module, which is designed to solve the problem of no information exchange and no feature interaction in a two-stream network, and is mainly used for cross-fusion between two-stream networks. By introducing multiple convolutional layers, batch normalization layers, ReLU activation function modules and channel attention modules, the cross-fusion module achieves efficient fusion of information between channels and feature enhancement, thereby improving the overall performance of the quality assessment model.

[0056] Specifically, the cross-fusion module includes a first channel, a second channel, a first cross-fusion addition module, a second cross-fusion addition module, a first SE module, a second SE module, a first post-processing module and a second post-processing module. The first channel includes a first cross-fusion convolution layer, a first cross-fusion batch normalization layer and a first cross-fusion ReLU activation function module, and the second channel includes: a second cross-fusion convolution layer, a second cross-fusion batch normalization layer and a second cross-fusion ReLU activation function module. The first post-processing module includes: a third cross-fusion convolution layer, a third cross-fusion batch normalization layer and a third cross-fusion ReLU activation function module; the second post-processing module includes: a fourth cross-fusion convolution layer, a fourth cross-fusion batch normalization layer and a fourth cross-fusion ReLU activation function module; In this embodiment, the first cross-fusion convolution layer and the second cross-fusion convolution layer are both 3 3 convolutional layers (Conv 3 3) is used to adjust the number of feature channels so that the output of the first channel and the output of the second channel have the same size.

[0057] The inputs of the first channel and the second channel come from the first output of the first network and the first output of the second network of the dual-stream network module, respectively. First, the input of the first channel is processed by the first cross-fusion convolution layer, the first cross-fusion batch normalization layer, and the first cross-fusion ReLU activation function module in sequence to extract local features; at the same time, the input of the second channel is processed by the second cross-fusion convolution layer, the second cross-fusion batch normalization layer, and the second cross-fusion ReLU activation function module in sequence to extract local features. The output of the first channel and the first output of the second network are added in the first cross-fusion addition module, and the output of the second channel and the first output of the first network are added in the second cross-fusion addition module. Then, the features are enhanced through two SE modules (the first SE module and the second SE module), where the first SE module performs feature selection on the output of the first cross-fusion addition module; and the second SE module performs feature selection on the output of the second cross-fusion addition module. Finally, the output of the first SE module (i.e., the enhanced features) is processed by the third cross-fusion convolution layer, the third cross-fusion batch normalization layer, and the third cross-fusion ReLU activation function module in turn to obtain the first CFM output; the output of the second SE module (i.e., the enhanced features) is processed by the fourth cross-fusion convolution layer, the fourth cross-fusion batch normalization layer, and the fourth cross-fusion ReLU activation function module in turn to obtain the second CFM output to achieve dimensionality reduction and nonlinear transformation, thereby reducing the number of feature channels and extracting higher-level features. After these processes, two sets of features that have been enhanced and fused are output. The two sets of features belong to high-quality classifiers and low-quality classifiers, respectively, and serve as the input of the subsequent residual channel attention module. This cross-fusion module significantly improves the performance and efficiency of neural networks through inter-channel information fusion and feature enhancement, and is suitable for various deep learning tasks such as image classification, object detection, semantic segmentation, etc.

[0058] In this embodiment, the third cross-fusion convolution layer and the fourth cross-fusion convolution layer are both 1 1 convolutional layer (Conv 1 1) is used to adjust the number of feature channels so that the size of the first CFM output and the second CFM output are the same.

[0059] (3) Residual channel attention module The present invention provides a residual channel attention module, which is designed to use residual connections to avoid the problem of gradient vanishing during the training of a quality assessment model, and at the same time use the attention mechanism to complete feature selection and generate final classification features. The residual channel attention module is designed to achieve communication and interaction between deep features (i.e., the second output of the first network and the second output of the second network) and multi-scale features (i.e., the first CFM output and the second CFM output). Through multiple convolutional layers, global pooling layers, ReLU activation functions and Sigmoid activation functions, the residual channel attention module effectively integrates deep features and multi-scale features, and enhances the feature expression ability of the quality assessment model.

[0060] Specifically, two RCAM modules are set, the first RCAM module is set corresponding to the high-quality classifier, and the second RCAM module is set corresponding to the low-quality classifier. Each RCAM module has three parallel branches, and the RCAM module has two inputs: one is the output from the corresponding CFM module, and the other is the output from the last dense connection block of the DenseNet network (the first network or the second network). The first RCAM module includes: the first residual convolution layer, the first residual addition module, the second residual convolution layer, the first Sigmoid activation function module, the third residual convolution layer, the first multiplication module, the first global pooling layer, the fourth residual convolution layer, the first residual ReLU activation function module, and the second multiplication module; the second RCAM module includes: the fifth residual convolution layer, the second residual addition module, the sixth residual convolution layer, the second Sigmoid activation function module, the seventh residual convolution layer, the third multiplication module, the second global pooling layer, the eighth residual convolution layer, the second residual ReLU activation function module, and the fourth multiplication module.

[0061] For the first RCAM module, in the first branch, the first CFM output is adjusted through the first residual convolution layer and the second residual convolution layer to adjust the number of feature channels, wherein the output of the first residual convolution layer and the second output of the first network are added at the feature element level in the first residual addition module, and after being processed by the second residual convolution layer, the first attention weight is generated by the first Sigmoid activation function module. In the second branch, the first CFM output is convolved by the third residual convolution layer. Subsequently, the first attention weight is multiplied by the output of the third residual convolution layer in the first multiplication module to achieve feature selection and obtain the first selected feature. In the third branch, the first CFM output is processed by the first global pooling layer, the fourth residual convolution layer and the first residual ReLU activation function module to obtain the first channel attention score. Finally, the above-mentioned first selected feature and the first channel attention score are multiplied in the second multiplication module to obtain the first classification feature.

[0062] For the second RCAM module, in the first branch, the second CFM output is adjusted through the fifth residual convolution layer and the sixth residual convolution layer to adjust the number of feature channels, wherein the output of the fifth residual convolution layer and the second output of the second network are added at the feature element level in the second residual addition module, and after being processed by the sixth residual convolution layer, the second attention weight is generated by the second Sigmoid activation function module. In the second branch, the second CFM output is convolved by the seventh residual convolution layer. Subsequently, the second attention weight is multiplied by the output of the seventh residual convolution layer in the third multiplication module to realize feature selection and obtain the second selected feature. In the third branch, the second CFM output is processed by the second global pooling layer, the eighth residual convolution layer and the second residual ReLU activation function module to obtain the second channel attention score. Finally, the above second selected feature is multiplied by the second channel attention score in the fourth multiplication module to obtain the second classification feature.

[0063] The residual channel attention module effectively avoids the gradient vanishing problem and enhances the feature selection capability through residual connection and channel attention mechanism. It is suitable for various deep learning tasks such as image classification, object detection, semantic segmentation, etc.

[0064] In this embodiment, the first residual convolution layer, the second residual convolution layer, the third residual convolution layer, the fifth residual convolution layer, the sixth residual convolution layer and the seventh residual convolution layer are all 3 3 convolutional layers, the fourth residual convolutional layer and the eighth residual convolutional layer are 1 1 convolutional layer.

[0065] (4) Classification module The classification module of the present invention specifically includes: A third global pooling layer is used to perform global pooling processing on the first classification feature; A fourth global pooling layer, used for performing global pooling processing on the second classification feature; The first fully connected layer maps the first classification feature of the global pooling to the high-quality Bernoulli distribution classifier to obtain the first Bernoulli probability; The second fully connected layer is used to map the globally pooled second classification features to the low-quality Bernoulli distribution classifier to obtain the second Bernoulli probability; The quality category judgment module is used to obtain the quality category according to the first Bernoulli probability and the second Bernoulli probability.

[0066] The present invention uses a fully connected layer to map the extracted first classification features and second classification features to Bernoulli distributions of different quality categories, and obtains Bernoulli probabilities corresponding to high-quality classifiers and low-quality classifiers, respectively. The Bernoulli probability corresponding to the high-quality classifier is a value between 0 and 1, which is intended to determine whether the fundus image is a high-quality image. The Bernoulli probability corresponding to the low-quality classifier is a value between 0 and 1, which is intended to determine whether the fundus image is a low-quality image.

[0067] The fundus image quality classification method is as follows: if the Bernoulli probability is less than 0.5, it is set to 0, and if it is greater than 0.5, it is set to 1; if the high-quality classifier and the low-quality classifier obtain (1, 0) respectively, the predicted image quality level is 0; if the high-quality classifier and the low-quality classifier obtain (0, 0) respectively, the predicted image quality level is 1; if the high-quality classifier and the low-quality classifier obtain (0, 1) respectively, the predicted image quality level is 2. Quality level 1 is intermediate quality, and this method can effectively assist in the classification of intermediate quality images. Through the design of a dual-stream network, the present invention improves the classification accuracy of intermediate quality images. After the intermediate quality fundus images are separated, they can be subsequently enhanced, and then doctors use them for clinical diagnosis of patients, which greatly improves the efficiency of fundus image use.

[0068] (5) GradCAM visualization processing module The present invention utilizes GradCAM technology to perform visual analysis on degradation features and structural features, and provides an interpretable basis for fundus image feature distribution for fundus image quality classification results.

[0069] The GradCAM visualization processing module of the present invention uses GradCAM technology to visualize the areas in the fundus image that affect the classification decision. The GradCAM visualization processing module generates a heat map by calculating the gradient of the first classification feature relative to the output of the third residual convolution layer and the gradient of the second classification feature relative to the output of the seventh residual convolution layer, highlighting the image area that the quality assessment model pays attention to when making decisions. This visualization method not only displays the degradation features (such as blur, noise) and structural features (such as vascular structure and fundus pathological structure) that the quality assessment model pays attention to, but also helps doctors understand the classification basis of the quality assessment model and improves the credibility of the classification results.

[0070] By combining the Bernoulli distribution output of the fully connected layer and the GradCAM visualization analysis, the present invention achieves high-accuracy fundus image quality classification and provides an intuitive feature distribution interpretation. It is suitable for the field of medical image analysis, especially the quality classification task of fundus images, and can provide effective auxiliary support in practical applications. It has broad application prospects and significant practical value.

[0071] Example 2 This embodiment provides a fundus image quality assessment method, which is implemented based on the fundus image quality assessment system provided in Embodiment 1, and includes: Acquiring fundus images to be evaluated; Inputting the fundus image into a pre-trained quality assessment model for quality classification to obtain a quality category; Outputs the quality category.

[0072] The present invention conducted quality assessment experiments on a batch of fundus images of different qualities, and the results are shown in Table 1.

[0073]

[0074] Table 1 shows the experimental results of different methods in the fundus image quality assessment task, comparing the accuracy, AUC area, F1 score and Kappa score. Compared with Vision Transformer (ViT), DenseNet, Inception-V3, MCF-Net (Multiple Color-space Fusion Network) and FIQuA (Fundus Image Quality Assessment) (Aditya Raj, Multivariate Regression-Based Convolutional Neural Network Model for Fundus Image Quality Assessment, IEEE Access, 2020, (8): 57810-57821), the proposed method performs well in all indicators and achieves the best performance: the accuracy is 87.81%, the AUC area is 85.45%, the F1 score is 79.49%, and the Kappa score is 72.10%. Compared with the second-best methods FIQuA (83.31%) and Inception-V3 (82.95%) in terms of accuracy, the proposed method significantly improves the evaluation performance. This is mainly due to the introduction of the cross-fusion module and the residual channel attention module, which makes the distribution of features at different levels more separated, especially improves the distinction between level 0 and level 1, thereby effectively reducing confusion and ensuring higher classification accuracy and stability.

[0075] In another embodiment of the present invention, a computer device is provided, the computer device includes a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the fundus image quality assessment method.

[0076] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed random access memory (Random Access Memory, RAM) or a non-volatile memory (non-volatile memory), such as at least one disk storage. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the fundus image quality assessment method in the above embodiment.

[0077] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0078] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0079] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A fundus image quality assessment system, characterized in that: include: A quality assessment model is used to classify the quality of fundus images and obtain a quality category; The quality assessment model includes: A dual-stream network module, comprising a first network and a second network for respectively extracting high-quality image-related features and low-quality image-related features in fundus images; the first network and the second network both comprise a plurality of densely connected blocks connected in sequence; the outputs of the densely connected blocks of the first network are concatenated as the first output of the first network, and the output of the last densely connected block is the second output of the first network; the outputs of the densely connected blocks of the second network are concatenated as the first output of the second network, and the output of the last densely connected block is the second output of the second network; A cross-fusion module, used for performing cross-fusion processing on the first output of the first network and the first output of the second network to obtain a first CFM output and a second CFM output; A first RCAM module is used to focus attention on the first CFM output and the second output of the first network to obtain a first classification feature; A second RCAM module; used for focusing attention on the second CFM output and the second network second output to obtain a second classification feature; The classification module is used to obtain a quality category according to the first classification feature and the second classification feature.

2. The fundus image quality assessment system according to claim 1, characterized in that: It also includes a Gabor filtering module; the Gabor filtering module is used to perform Gabor filtering on the fundus image to obtain a Gabor filter image; the input image of the first network is the fundus image, and the input image of the second network is the Gabor filter image.

3. The fundus image quality assessment system according to claim 1, characterized in that: The cross-fusion module includes: A first channel is used to extract local features of a first output of the first network; A first cross-fusion summing module, used for adding the output of the first channel to the first output of the second network; A first SE module, used for performing feature selection on the output of the first cross-fusion summation module; A first post-processing module, used for adjusting the number of characteristic channels of the output of the first SE module so that the number of characteristic channels of the output of the first SE module is consistent with the number of characteristic channels of the first output of the first network, and obtaining a first CFM output; A second channel is used to extract local features of the first output of the second network; A second cross-fusion summing module is used to perform addition processing on the output of the second channel and the first output of the first network; The second SE module is used to perform feature selection on the output of the second cross-fusion summation module; The second post-processing module is used to adjust the number of feature channels of the output of the second SE module so that the number of feature channels of the output of the second SE module is consistent with the number of feature channels of the first output of the second network, thereby obtaining a second CFM output.

4. The fundus image quality assessment system according to claim 3, characterized in that: The first channel includes: a first cross-fusion convolution layer, a first cross-fusion batch normalization layer and a first cross-fusion ReLU activation function module; the second channel includes: a second cross-fusion convolution layer, a second cross-fusion batch normalization layer and a second cross-fusion ReLU activation function module; The first output of the first network is processed in sequence by a first cross-fusion convolutional layer, a first cross-fusion batch normalization layer, and a first cross-fusion ReLU activation function module to obtain an output of a first channel; The first output of the second network is processed in sequence by a second cross-fusion convolutional layer, a second cross-fusion batch normalization layer, and a second cross-fusion ReLU activation function module to obtain an output of a second channel.

5. The fundus image quality assessment system according to claim 1, characterized in that: The first RCAM module comprises: A first residual convolution layer, used to perform a convolution operation on the first CFM output; A first residual summing module, used for performing feature element-level addition of the output of the first residual convolutional layer and the second output of the first network; A second residual convolution layer, used to perform a convolution operation on the output of the first residual summation module; A first Sigmoid activation function module is used to process the output of the second residual convolution layer using the Sigmoid activation function to obtain a first attention weight; The third residual convolution layer is used to perform a convolution operation on the first CFM output; A first multiplication module, used for multiplying the first attention weight and the output of the third residual convolution layer to obtain a first selected feature; A first global pooling layer, used for performing pooling processing on the first CFM output; A fourth residual convolution layer, used to perform a convolution operation on the output of the first global pooling layer; A first residual ReLU activation function module, used to process the output of the fourth residual convolution layer using a ReLU activation function to obtain a first channel attention score; The second multiplication module is used to multiply the first selected feature by the first channel attention score to obtain a first classification feature.

6. The fundus image quality assessment system according to claim 5, characterized in that: The second RCAM module comprises: A fifth residual convolution layer, used to perform a convolution operation on the second CFM output; A second residual summing module, used for performing feature element-level addition of the output of the fifth residual convolutional layer and the second output of the second network; A sixth residual convolution layer, used for performing a convolution operation on the output of the second residual summation module; A second Sigmoid activation function module is used to process the output of the sixth residual convolution layer using a Sigmoid activation function to obtain a second attention weight; A seventh residual convolution layer, used for performing a convolution operation on the second CFM output; A third multiplication module is used to multiply the second attention weight and the output of the seventh residual convolution layer to obtain a second selected feature; A second global pooling layer, used for performing pooling processing on the second CFM output; An eighth residual convolution layer, used to perform a convolution operation on the output of the second global pooling layer; A second residual ReLU activation function module, used to process the output of the eighth residual convolution layer using the ReLU activation function to obtain a second channel attention score; The fourth multiplication module is used to multiply the second selected feature by the second channel attention score to obtain a second classification feature.

7. The fundus image quality assessment system according to claim 1, characterized in that: It also includes a data augmentation module; The data augmentation module is used to expand the fundus images of a few quality categories in the training data set before training the quality assessment model.

8. A method for evaluating fundus image quality, characterized in that: The fundus image quality assessment system according to any one of claims 1 to 7 is used to classify the quality of the fundus image to obtain a quality category.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the fundus image quality assessment method according to claim 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the fundus image quality assessment method according to claim 8 is implemented.

Citation Information

Patent Citations

  • Image no-reference quality evaluation method and device based on visual interactive perception double-flow network

    CN115457015A

  • Eye fundus image recognition method and device based on binocular feature fusion and storage medium

    CN117912092A

  • Image Quality Assessment Using Similar Scenes as Reference

    US20190122115A1

  • Device and method for glaucoma auxiliary diagnosis, and storage medium

    US20220130052A1