Eye fundus graph quality evaluation network and evaluation method combining global and local features

By combining global and local features, the fundus image quality evaluation network is used to enhance the G channel and cross attention mechanism of RGB images, the problem of inaccurate fundus image quality evaluation in the prior art is solved, more efficient feature extraction and fusion is achieved, and the accuracy and stability of fundus image quality evaluation is improved.

CN120339191APending Publication Date: 2025-07-18HUNAN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510321536.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing fundus image quality assessment methods cannot fully reflect the clinical diagnostic value of retinal images, especially in low-quality images that are susceptible to blur, noise or light changes, resulting in inaccurate evaluation.

Method used

A fundus graphics quality evaluation network combining global and local features is adopted, and the fundus graphics quality evaluation network is used to enhance the G channel of the RGB image by independent global feature extraction branches and local feature extraction branches, and the G-channel of RGB images is enhanced, combining cross attention mechanisms and multiple weighted loss strategies to feature fusion to highlight key features in fundus images.

Benefits of technology

It improves the accuracy and stability of fundus image quality evaluation, can more effectively capture the interactive information of global and local features, and improves the performance of image quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339191A_ABST
    Figure CN120339191A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fundus image quality evaluation, and particularly discloses a fundus image quality evaluation network and evaluation method combining global and local features, and the network comprises a feature extraction module which comprises a global feature extraction branch and a local feature extraction branch which are independently operated, applying an RGB image enhanced by CLAHE to a G channel, and inputting a local feature extraction branch; the feature fusion module is used for fusing the global features and the local features output by the feature extraction module by adopting a cross attention mechanism and a plurality of full connection layers, and respectively providing independent supervision signals for a global feature extraction branch and a local feature extraction branch through a three-way weighting loss strategy; according to the method, key features in the fundus image are effectively highlighted, deep fusion of global and local features is realized, richer feature interaction information is captured, and the method has excellent performance in quality evaluation of the fundus image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fundus image quality assessment, and specifically discloses a fundus image quality assessment network and method that combines global and local features. Background Art

[0002] Traditional retinal image quality assessment methods can be divided into two categories. One is to evaluate based on general quality parameters (such as sharpness, contrast, illumination uniformity, etc.). For example, the quality evaluation algorithm focuses on the sharpness and illumination of the image, and measures the illumination quality by evaluating the contrast and brightness of the retinal image. They divide the image into non-overlapping square regions, analyze these regions separately, and then combine the quality indicators of these regions to form an overall quality indicator. However, the features extracted above are all low-level features and cannot comprehensively reflect the clinical diagnostic value of retinal images, especially cannot directly measure the sharpness or integrity of anatomical structures. At the same time, analyzing the image after dividing it into independent regions may ignore the integrity and relevance of the retinal structure, and when encountering some specific quality problems (such as local blur, glare or occlusion), it will lead to inaccurate overall quality assessment. The other is based on structural quality parameters, such as the visibility of the macula, anatomical structures such as the optic disc and blood vessels. For example, the image quality is judged based on the fundus structure characteristics. First, the blood vessels are segmented by a directional matching filter and region combination segmentation method, and then the sharpness and area of each segment of blood vessels are used as quality evaluation indicators. However, this method requires precise segmentation of the retinal structure (such as blood vessels, optic disc, etc.) first, but in images with poor quality, the segmentation may be affected by blur, noise or illumination changes, thus reducing the reliability of the assessment.

[0003] In recent years, convolutional neural networks have achieved great success in image classification tasks. Retinal image quality assessment is essentially also an image classification task, so convolutional neural networks have recently been widely introduced into the retinal image quality assessment task. For example, by combining unsupervised saliency maps and supervised CNN feature extraction, saliency region detection mainly relies on the model to segment and identify important regions in the image. However, this method has limited effectiveness in identifying important structures in complex scenes or low-quality images. If the key regions in the image (such as the optic disc or blood vessels) are not correctly detected, the subsequent quality assessment will be greatly affected. Some researchers have also proposed a multi-color space fusion network (MCF-Net) for retinal image quality assessment. MCF-Net unifies three parallel CNN branches into a framework to learn complementary information contexts for retinal IQA from the RGB, HSV, and Lab color spaces. Although RGB, HSV, and Lab represent different color characteristics, there is still some redundant information between them. For example, although HSV and Lab are different color spaces, they both have separate luminance channels, which may cause some of the information captured by the network to be repeated when extracting features, thus reducing the efficiency of the model. In addition, MCF-Net extracts features in different color spaces through three parallel CNN branches, which greatly increases the number of model parameters, complexity, and computational cost. Some researchers have also used a saliency structure detector to obtain saliency maps of anatomical structures, concatenate them, and then input them into a CNN. Similar to the method based on structural quality parameters, when the overall quality of the image is poor, it will be difficult to clearly extract the structure, thus affecting the evaluation results. Some researchers have also proposed a dark and bright channel prior-guided deep network for retinal image quality assessment. Without adding additional parameters, it introduces the dark channel and bright channel priors into the deep network and allows end-to-end training. The dark and bright channel prior assumes that it holds for all retinal images, but this prior may fail in cases of uneven illumination, severe glare, or high noise. For example, overexposed regions may cause the loss of dark channel information, while low-luminance regions may make the bright channel unreliable.

[0004] Therefore, in view of this, the inventors have provided a fundus image quality assessment network and assessment method that combines global and local features to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a fundus image quality assessment network that combines global and local features and has excellent and stable performance to improve the current fundus image quality assessment method.

[0006] To achieve the above purpose, the basic solution of the present invention provides a fundus image quality assessment network that combines global and local features, including:

[0007] Feature extraction module: It includes a global feature extraction branch and a local feature extraction branch that operate independently. The input of the global feature extraction branch is an RGB image, and the output is the global feature of the RGB image. The input of the local feature extraction branch is the RGB image after applying CLAHE enhancement to the G channel, and the output is the local feature of the RGB image;

[0008] Feature fusion module: It uses a cross-attention mechanism and multiple fully connected layers to fuse the global feature and local feature output by the feature extraction module, and provides independent supervision signals for the global feature extraction branch and the local feature extraction branch through a three-way weighted loss strategy.

[0009] Furthermore, the global feature extraction branch includes multiple stages. Each stage includes DownSample and ConvNeXtBlock, and a channel attention module is added after each stage. After global average pooling and layer normalization of the features output by the last stage, the global feature of the RGB image is obtained;

[0010] The local feature extraction branch includes multiple DenseBlocks and Transition layers, and a spatial attention module is added after each DenseBlock. After global average pooling and layer normalization of the features output by the last Transition layer, the local feature of the RGB image is obtained.

[0011] Furthermore, in the three-way weighted loss strategy:

[0012] The global feature and local feature output by the feature extraction module are input into a linear classifier for classification, and two loss values are obtained respectively;

[0013] The global feature and local feature pass through a bidirectional cross-attention block, and after concatenation, feature dimensionality reduction is performed through multiple fully connected layers, and then input into the classifier for classification again, and another loss value is obtained;

[0014] The three loss values are weighted to obtain the final loss, in order to capture richer feature interaction information.

[0015] Furthermore, the process of applying CLAHE enhancement to the G channel of the RGB image is as follows:

[0016] The G channel of the RGB image is divided into multiple grids that can be processed independently;

[0017] Calculate the gray-level histogram of each grid;

[0018] Set a clipping threshold and limit the frequency of each gray level to clip the histogram;

[0019] Evenly distribute the cropped frequency to all gray levels;

[0020] Calculate the cumulative distribution function according to the cropped histogram;

[0021] Map the cumulative distribution function to the target gray value range through normalization, which is the enhanced result of each grid;

[0022] For the boundaries between adjacent grids, apply bilinear interpolation to ensure the smoothness of the transition region, and combine the enhanced G channel with the R channel and B channel to obtain the RGB image with the G channel enhanced by CLAHE.

[0023] Furthermore, in the cross-attention mechanism:

[0024] Perform linear transformations on the global feature and the local feature to generate a query vector Q, a key vector K, and a value vector V respectively;

[0025] Calculate the similarity between the query vector Q and the key vector K through dot product to obtain the attention score;

[0026] Weighted sum the value vector V with the attention score;

[0027] Combine the result of the weighted sum with the global feature or the local feature through a residual connection method to obtain the combined global feature or local feature;

[0028] Concatenate the combined global feature or local feature to obtain the final feature, and input the final feature into the fully connected layer and finally use it for the classification task.

[0029] Furthermore, in the process of obtaining the combined global feature, the global feature extracted by the global feature extraction branch is used as the query vector Q, and the local feature extracted by the local feature extraction branch is used as the key vector K and the value vector V;

[0030] In the process of obtaining the combined local feature, the local feature extracted by the local feature extraction branch is used as the query vector Q, and the global feature extracted by the global feature extraction branch is used as the key vector K and the value vector V.

[0031] Based on the same inventive concept, the present invention discloses a method for evaluating the quality of fundus images by combining global and local features, including using the above-mentioned fundus image quality evaluation network to process the fundus images and output the processed images for evaluation and classification.

[0032] Furthermore, the steps of using the above-mentioned fundus image quality evaluation network to process the fundus images and output the processed images for evaluation and classification are as follows:

[0033] Step S1: Obtain the original retinal RGB image to be evaluated;

[0034] Step S2: Use the feature extraction module in the above fundus image quality assessment network to obtain the global features and local features in the original retinal RGB image respectively;

[0035] Step S3: Use the feature fusion module in the above fundus image quality assessment network to fuse the obtained global features and local features, and obtain the final fused features;

[0036] Step S4: Input the final features into the fully connected layer and use them for the evaluation and classification tasks of fundus images.

[0037] Furthermore, in step S2, when applying CLAHE to enhance the local features of the G channel, apply CLAHE enhancement to the G channel of the original retinal RGB image.

[0038] Furthermore, in step S3, provide independent supervision signals for the global feature extraction branch and the local feature extraction branch respectively through a three-way weighted loss strategy.

[0039] The principle and effect of this solution are as follows:

[0040] 1. Different from the traditional method of converting the RGB image to the LAB image and then applying CLAHE to the L channel, the present invention directly applies CLAHE enhancement to the G channel of the RGB image. This can clearly display the key structural features in the fundus image, such as retinal blood vessels, macular areas, and diseased tissues, thereby improving the ability to extract local features. The L channel in the LAB color space mainly reflects the brightness information of the image. Applying CLAHE to it usually can enhance the overall brightness and contrast, but the effect of enhancing local structural details may be limited. In the G channel, blood vessels have a higher contrast compared to the background visually and contain rich local blood vessel structure information. Therefore, applying CLAHE to the G channel can more effectively highlight the key features in the fundus image.

[0041] 2. The three-way weighted loss strategy can provide independent supervision signals for the global feature extraction branch and the local feature extraction branch respectively, ensuring that the feature extraction ability of each branch is optimal, avoiding interference between branches, enabling the global feature extraction branch to focus on improving the classification ability of global information, and the local feature extraction branch to focus on improving the classification ability of local details. The fused loss can then achieve the deep fusion of the global and local features extracted by the two branches, capturing richer feature interaction information.

[0042] 3. By applying CLAHE enhancement to the G channel in the RGB image, the key features in the fundus image are effectively highlighted, and a three-way weighted loss strategy is used to achieve deep fusion of global and local features, capturing richer feature interaction information, which has superior performance in the quality assessment of fundus graphics. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1 FIG. shows a schematic diagram of a fundus graphic quality assessment network combining global and local features proposed in an embodiment of the present application;

[0045] Figure 2 FIG. shows a schematic diagram of applying CLAHE enhancement to the G channel in a fundus graphic quality assessment network combining global and local features proposed in an embodiment of the present application;

[0046] Figure 3 FIG. shows a schematic diagram of the cross-attention mechanism in a fundus graphic quality assessment network combining global and local features proposed in an embodiment of the present application;

[0047] Figure 4 FIG. shows a comparison diagram of the confusion matrix diagrams of a fundus graphic quality assessment network combining global and local features proposed in an embodiment of the present application and the prior art tested on the EyeQ dataset, where (a) is the confusion matrix diagram of GuidedNet tested on the EyeQ dataset, (b) is the confusion matrix diagram of SalStructIQA tested on the EyeQ dataset; (c) is the confusion matrix diagram of DualRegNet tested on the EyeQ dataset; (d) is the confusion matrix diagram of DGL-Net proposed in an embodiment of the present application tested on the EyeQ dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and effects of the present invention as follows.

[0049] A fundus graphic quality assessment network combining global and local features, as shown in the embodiments Figure 1 below: It includes a feature extraction module and a feature fusion module. Among them, the feature extraction module further includes a global feature extraction branch and a local feature extraction branch.Figure 1 The explanations of each character are as follows: C for splicing, CA for cross attention, Spatial Attention for spatial attention, Conv for global feature extraction, Loss for loss, X for element-wise multiplication, Channel Attention for channel attention, and Dense for local feature extraction. The fundus image quality assessment network is specifically as follows:

[0050] The RGB image is the original retinal image collected. The RGB image is used as the input of the global feature extraction branch, and after applying CLAHE enhancement to the G channel in the RGB image, it is used as the input of the local feature extraction branch.

[0051] In this embodiment, the original RGB image is represented by I = (R, G, B). After applying ClAHE enhancement to the G channel, it can be expressed as:

[0052] G GLAHE = GLAHE(G)

[0053] The green channel after CLAHE enhancement is combined with other channels to form a new RGB image I' as the input:

[0054] I' = (R, G CLAHE , B)

[0055] I' is the RGB image after applying CLAHE enhancement to the G channel.

[0056] In the feature extraction module, the global feature extraction branch and the local feature extraction branch operate independently, and a channel attention mechanism or a spatial attention mechanism is added after each stage for different branches.

[0057] For example, the global feature extraction branch consists of multiple stages. Each stage is composed of a downsampling layer DownSample and a ConvNeXtBlock, and a channel attention module represented by ECA(·) is added after each stage. In this embodiment, assuming the input image is I, the process of the global feature extraction branch can be expressed as:

[0058] F0 = I

[0059] F i = ECA(ConvNeXtBlock i (Down i (F i-1 ))), i = 1, 2,..., N

[0060] Among them, F i represents the feature output at the i-th stage. Finally, a global average pooling GAP and a layer normalization LN operation are performed on the feature F N output at the N-th stage to obtain the final global feature:

[0061] F global = LN(GAP(F N ))

[0062] F global is the final global feature.

[0063] The local feature extraction branch consists of multiple DenseBlocks and Transition layers. A spatial attention module, denoted as SA(·), is added after each DenseBlock. If the input is I‘, the process of local feature extraction can be expressed as:

[0064] F0 = Stem(I')

[0065] F i = SA(DenseBlock i (F i-1 ), i = 1, 2, … N

[0066] F i+1 = Transition i (F i ), i = 1, 2, … N - 1

[0067] where Stem(·) represents the initial convolutional layer and pooling operation of the local feature extraction branch. Finally, batch normalization operation (BN) and global average pooling operation (GAP) are performed on F N to obtain the final local feature:

[0068] F local = GAP(BN(F N ))

[0069] F local is the final local feature.

[0070] In the feature fusion module, a three-way weighted loss strategy is adopted. The feature vectors finally output by each branch are input into a linear classifier for classification to obtain the loss value. The expression is as follows:

[0071]

[0072]

[0073] where y is the true label, and are the predicted values of their respective branches.

[0074] After that, the feature vectors output from each branch pass through a bidirectional cross-attention block, denoted as DCA(·). Then, the two feature vectors are concatenated and passed through multiple fully connected layers for feature dimensionality reduction. Finally, they are input into a classifier for classification, and a loss value is obtained again. The expression is as follows:

[0075] F' global = DCA(F global , F local , F local )

[0076] F' local = DCA(F local , F global , F global )

[0077] F fusion = Concat(F' global , F' local )

[0078] F reduced = FC i (F fusion )(i = 1, 2, …N)

[0079]

[0080] The final loss is obtained by weighting three losses:

[0081] L total = λ global L global + λ locak L local + λ fusion L fusion

[0082] where λ global , λ local , λ fusion are the corresponding weight hyperparameters.

[0083] The three-way weighted loss strategy can provide independent supervision signals for the ConvNeXt and DenseNet branches respectively, ensuring the optimal feature extraction ability of each branch, avoiding interference between branches, enabling the ConvNeXt branch to focus on improving the classification ability of global information, and the DenseNet branch to focus on improving the classification ability of local details. The fused loss can achieve a deep fusion of the global and local features extracted by the two branches, capturing richer feature interaction information.

[0084] Different from applying CLAHE to the L channel after converting an RGB image to an LAB image using traditional methods, in this embodiment, CLAHE enhancement is directly applied to the G channel of the RGB image to clearly display key structural features in fundus images, such as retinal blood vessels, macular regions, and diseased tissues, thereby improving the ability to extract local features. The L channel in the LAB color space mainly reflects the brightness information of the image. Applying CLAHE to it usually can enhance the overall brightness and contrast, but may have limited effects on enhancing local structural details. In the G channel, blood vessels have a higher contrast visually compared to the background and contain rich local blood vessel structure information. Therefore, applying CLAHE to the G channel can more effectively highlight the key features in fundus images. Based on this, in this embodiment, CLAHE enhancement is applied to the G channel. The whole process is as Figure 2 shown as follows:

[0085] First, extract the G channel from an RGB image of size H×W, then divide the G channel into grids of size m×n, and then process each grid independently to ensure that the enhancement effect adapts to the local area. The number of grids is:

[0086]

[0087] Calculate the gray-level histogram H(g) for each grid, where g∈[0,L - 1] is the gray level, L is usually 256, and H(g) is the number of pixels at gray level g. The calculation formula is as follows:

[0088]

[0089] where I(i,j) is the image pixel value, and δ(*) is the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise.

[0090] After calculating the histogram for each grid, the histogram needs to be clipped. Set the clipping threshold as T to limit the frequency H(g) of each gray level to avoid excessive contrast enhancement. The clipping threshold T is calculated by the following formula:

[0091]

[0092] where C is the contrast limiting coefficient (a hyperparameter, usually 2 - 4), and M×N is the total number of pixels in each small block. If the pixel frequency of a certain gray level exceeds T, the excess part is evenly distributed to other gray levels:

[0093]

[0094] Distribute the clipped frequency evenly to all gray levels to maintain the total number of pixels. The expression is as follows:

[0095]

[0096] Then calculate the cumulative distribution function (CDF). According to the cropped histogram H″(g), calculate the cumulative distribution function C(g) for gray value mapping. The expression is as follows:

[0097]

[0098] After that, perform normalized CDF and map C(g) to the target gray value range (usually [0, 255]):

[0099]

[0100] where C min is the smallest non - zero cumulative frequency, and C min is the largest cumulative frequency. After mapping the gray value of each pixel g in the grid to the enhanced value g′ through the above formula, the enhanced result of each grid is obtained.

[0101] Finally, for the boundaries between adjacent grids, apply bilinear interpolation to ensure the smoothness of the transition region. Then combine the enhanced G channel with the R channel and B channel to obtain the RGB image with the G channel enhanced by CLAHE.

[0102] The global feature extraction branch in this embodiment is used to obtain global features that can reflect the overall quality of fundus images. The global feature extraction branch uses ConvNeXt as the backbone. ConvNeXt combines some key design concepts of Transformer on the basis of deep convolutional neural networks. Compared with traditional convolutional neural networks, ConvNeXt has a larger receptive field and uses layer normalization. These characteristics enable ConvNeXt to enhance its ability to capture global information while maintaining the efficiency of convolution. At the same time, ConvNeXt is also superior to the original Transformer. Although Transformer is very strong in global feature extraction, its computational cost is high and it has a strong dependence on large - scale data. As a convolutional network, ConvNeXt performs more stably and efficiently on medium - and small - scale datasets. Considering the dataset for fundus image quality assessment, using ConvNeXt is the most efficient. At the same time, in order to further improve the global feature extraction ability of the global feature extraction branch, a channel attention mechanism is added to the global feature extraction branch.

[0103] The local feature extraction branch in this embodiment is used to obtain local features (such as the fovea centralis, vascular structure, and peripapillary region) that reflect the local quality of fundus images. DenseNet is mainly used as the backbone. Compared with ConvNeXt, this network has smaller convolutional kernels and receptive fields, and is more suitable for capturing fine-grained features. At the same time, DenseNet uses dense connections, and the features of each layer are passed to all subsequent layers. This characteristic ensures that the local details (such as texture and edge information) extracted in the shallow layer will not be ignored by the deep network, which helps to retain key local information. At the same time, in order to further improve the extraction ability of local features of the local feature extraction branch, a spatial attention mechanism is added to the local feature extraction branch.

[0104] The feature fusion module in this embodiment is used to fuse global features and local features. The feature fusion module is mainly composed of a cross-attention mechanism and multiple fully connected layers.

[0105] Among them, the cross-attention mechanism is as Figure 3 shown. The descriptions of each character in Figure 3 are as follows: ⊕ and ⊙ represent element-wise addition and dot product respectively; Softmax is the activation function; AW is the weight matrix; Q, K, and V are the query vector, key vector, and value vector respectively. The input of the cross-attention mechanism is the one-dimensional feature vectors finally used for classification extracted by the global feature extraction branch and the local feature extraction branch, that is, F global and F local . For the global feature extraction branch, F global is used as the query vector Q, and F local is used as the key vector K and value vector V to better capture the supplementary information of the local branch to the global branch. Specifically, the calculation process of the cross-attention mechanism is as follows:

[0106] First, perform feature projection, and perform linear transformation on the input one-dimensional feature vectors F global and F local to generate the query vector Q, key vector K, and value vector V respectively. The formula is:

[0107] Q = W Q F global , K = W K F local , V = W V F local

[0108] where W Q , W K , W V are learnable weight matrices.

[0109] After that, calculate the similarity. Calculate the similarity between the query vector Q and the key vector K through the dot product to capture the correlation between the global feature and the local feature:

[0110]

[0111] where d k is the dimension of K for scaling to avoid overly large values. AW is the attention score.

[0112] After obtaining the attention score, the value vector V is weighted and summed using the attention score:

[0113] F ca = AW · V

[0114] where F ca is the feature after weighted summation.

[0115] Finally, to stabilize the training and retain the original information, the feature F ca and F global after cross-attention are combined through a residual connection:

[0116] F' global = F ca + F global

[0117] where F' global is the combined global feature.

[0118] For the local feature extraction branch, the same method is adopted, except that F local is used as Q, and F global is used as the key vector K and the value vector V, and finally the feature F' local of the combined local feature is obtained.

[0119] Through bidirectional cross-attention, the complementarity of global and local features can be fully utilized, enabling the two features to complement each other at different levels.

[0120] After that, the combined global feature F' global and the combined local feature F' local are concatenated to obtain the final feature F fusion , which is input into multiple fully connected layers to further extract high-order information and achieve non-linear mapping of features, and finally used for classification tasks.

[0121] To evaluate the reliability of the fundus image quality assessment network (DGL-Net) that combines global and local features proposed in this embodiment, extensive three-class classification tasks were performed using EyeQ in the experimental example. Quantitative analysis involves a detailed comparison between the fundus image quality assessment network that combines global and local features proposed in this embodiment and the prior art. The results of this evaluation are shown in Table 1, and the confusion matrix of the fundus image quality assessment network that combines global and local features proposed in this embodiment tested on the EyeQ dataset is as Figure 4 shown, which demonstrates the superior performance of the fundus image quality assessment network that combines global and local features proposed in this embodiment in the RIQA task.

[0122] Table 1 Comparison with the prior art on the EyeQ dataset

[0123]

[0124] This embodiment proposes technical methods including network structure, application of CLAHE enhancement to the G channel in RGB images, and multi-path loss. To illustrate the effectiveness of these methods, a series of ablation experiments were conducted on the EyeQ dataset in the experimental example.

[0125] As shown in Table 2, the influence of each component on the model was evaluated on the EyeQ dataset in the experimental example. Starting from local feature extraction, adding a global feature extraction branch, a feature fusion module, a channel & spatial attention mechanism, applying Clahe enhancement to the G channel, and finally forming the final DGL-Net model. After adding the global feature extraction branch, Kappa and Precision increased by 1.35% and 4.2% respectively. After adding channel and spatial attention, all indicators showed a significant improvement, indicating that the channel and spatial attention mechanism enhanced the capabilities of the global feature extraction branch and the local feature extraction branch. Finally, after enhancing the G channel in the experimental example and using it as the input of the local feature extraction branch, all indicators also had a good improvement.

[0126] Table 2 Results of ablation experiments on the EyeQ dataset

[0127]

[0128] To verify the effectiveness of applying CLAHE enhancement to the G channel in the RIQA task, a variety of comparative experiments were designed in the experimental example, and CLAHE enhancement was applied to the key channels of different color spaces respectively, including:

[0129] (1) The L channel in the LAB color space: The L channel mainly reflects the brightness information of the image. Applying CLAHE enhancement to it can improve the overall lighting conditions and brightness contrast, thereby highlighting the light and dark levels of the image.

[0130] (2) S and V channels in the HSV color space: The S channel represents saturation information. Enhancing its local contrast can improve the layering of color distribution; the V channel represents brightness, similar to the L channel in LAB. CLAHE enhancement helps improve brightness contrast and local details.

[0131] (3) Y channel in the YCbCr color space: The Y channel, as the luminance component, is similar to the L channel in LAB and the V channel in HSV. Applying CLAHE enhancement to it can improve brightness contrast, thereby enhancing the detail representation ability.

[0132] The experimental results are shown in Table 3. The experimental results indicate that when the image with CLAHE enhancement applied to the G channel is used as the input of the local feature extraction branch, the model achieves the optimal performance. This is because the G channel contains more vascular structure information than other channels in fundus images, and enhancing it can more significantly highlight local key features, thereby improving the classification effect of the RIQA task.

[0133] Table 3 Comparison of results for different input images on the EyeQ dataset

[0134]

[0135] When using multi-path losses, different experimental results will be obtained by assigning different weights to the losses of each branch. This is because different loss weights determine the influence degree of each loss term on the model during the optimization process. If the assignment is unreasonable, it may lead to the neglect of the training of some feature branches or fusion layers, thereby affecting the overall performance. The results of comparing the results of assigning different weight losses to each branch on the EyeQ dataset are shown in Table 4. Through experiments, it is found that when providing a loss weight of 0.2 for the global feature extraction branch and the local feature extraction branch respectively, the effect is the best, attributed to the fact that this weight assignment achieves the optimal balance of multi-branch collaborative training during the optimization process. The weight assignment of 0.2 ensures that the global and local feature branches play their respective roles during training, but does not overly emphasize the importance of a single branch, avoiding both excessive competition between branches and training instability or performance degradation caused by gradient conflicts.

[0136] Table 4 Comparison of results of assigning different weight losses to each branch on the EyeQ dataset

[0137]

[0138] In summary, the fundus image quality assessment network combining global and local features proposed in this embodiment effectively highlights the key features in the fundus image by applying CLAHE enhancement to the G channel in the RGB image, and uses a three-way weighted loss strategy to achieve deep fusion of global and local features, capturing richer feature interaction information, and has superior performance in the quality assessment of fundus images.

[0139] Based on the same inventive concept, an embodiment of the present invention discloses a method for combining global and local features for fundus image quality assessment, including using the above fundus image quality assessment network to process the fundus image and output the processed image for evaluation and classification. The steps are as follows:

[0140] Step S1, obtain the original retinal RGB image to be evaluated;

[0141] Step S2, use the global feature extraction branch in the feature extraction module of the fundus image quality assessment network to obtain the global features in the original retinal RGB image. After applying CLAHE enhancement to the G channel of the RGB image, use the local feature extraction branch in the feature extraction module of the fundus image quality assessment network to obtain the local features in the original retinal RGB image;

[0142] Step S3, use the feature fusion module in the fundus image quality assessment network to fuse the obtained global features and local features, and provide independent supervision signals for the global feature extraction branch and the local feature extraction branch respectively through a three-way weighted loss strategy, and obtain the final fused features;

[0143] Step S4, input the final features into the fully connected layer and use them for the evaluation and classification tasks of fundus images.

[0144] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes, but as long as the technical content of the present invention is not departed from, any indirect modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A fundus image quality assessment network combining global and local features, characterized in that Including: Feature extraction module: It includes a global feature extraction branch and a local feature extraction branch that operate independently. The input of the global feature extraction branch is an RGB image, and the output is the global feature of the RGB image. The input of the local feature extraction branch is the RGB image after applying CLAHE enhancement to the G channel, and the output is the local feature of the RGB image. Feature fusion module: It uses a cross-attention mechanism and multiple fully connected layers to fuse the global feature and local feature output by the feature extraction module, and provides independent supervision signals for the global feature extraction branch and the local feature extraction branch through a three-way weighted loss strategy.

2. The fundus image quality assessment network combining global and local features according to claim 1, characterized in that The global feature extraction branch includes multiple stages. Each stage includes DownSample and ConvNeXtBlock, and a channel attention module is added after each stage. The global feature of the RGB image is obtained after global average pooling and layer normalization of the features output by the last stage. The local feature extraction branch includes multiple DenseBlocks and Transition layers, and a spatial attention module is added after each DenseBlock. The local feature of the RGB image is obtained after global average pooling and layer normalization of the features output by the last Transition layer.

3. The fundus image quality assessment network combining global and local features according to claim 2, characterized in that, In the three-way weighted loss strategy: The global feature and local feature output by the feature extraction module are input into a linear classifier for classification, and two loss values are obtained respectively. The global feature and local feature pass through a bidirectional cross-attention block, and after concatenation, feature dimensionality reduction is performed through multiple fully connected layers, and then input into the classifier for classification again, and another loss value is obtained. The three loss values are weighted to obtain the final loss, so as to capture richer feature interaction information.

4. A fundus image quality assessment network combining global and local features according to claim 1, characterized in that The process of applying CLAHE enhancement to the G channel of the RGB image is as follows: The G channel of the RGB image is divided into multiple grids that can be processed independently. Calculate the gray-level histogram of each grid. Set the clipping threshold and limit the frequency of each gray level to clip the histogram. Evenly distribute the clipped frequency to all gray levels. Calculate the cumulative distribution function according to the clipped histogram. Map the cumulative distribution function to the target gray value range through normalization, which is the enhanced result of each grid. For the boundaries between adjacent grids, apply bilinear interpolation to ensure the smoothness of the transition region, and combine the enhanced G channel with the R channel and B channel to obtain the RGB image after applying CLAHE enhancement to the G channel.

5. The fundus image quality assessment network combining global and local features according to claim 1, characterized in that In the cross-attention mechanism: Perform linear transformation on the global feature and local feature to generate query vector Q, key vector K, and value vector V respectively. Calculate the similarity between query vector Q and key vector K through dot product to obtain attention scores. The attention scores are weighted and summed for value vector V. The result after weighted summation is combined with the global feature or local feature through a residual connection method to obtain the combined global feature or local feature. The combined global features or local features are concatenated to obtain the final features, and the final features are input into the fully connected layer and ultimately used for the classification task.

6. The fundus image quality assessment network combining global and local features according to claim 5, characterized in that, In the process of obtaining the combined global features, the global features extracted by the global feature extraction branch serve as the query vector Q, and the local features extracted by the local feature extraction branch serve as the key vector K and the value vector V. In the process of obtaining the combined local features, the local features extracted by the local feature extraction branch serve as the query vector Q, and the global features extracted by the global feature extraction branch serve as the key vector K and the value vector V.

7. A method for fundus image quality assessment by combining global and local features, characterized in that It includes using the fundus image quality assessment network according to any one of claims 1 to 6 to process the fundus image and output the processed image for evaluation and classification.

8. A method for fundus image quality assessment by combining global and local features as claimed in claim 7, characterized in that The steps of using the fundus image quality assessment network according to any one of claims 1 to 6 to process the fundus image and output the processed image for evaluation and classification are as follows: Step S1, obtain the original retinal RGB image to be evaluated. Step S2, use the feature extraction module in the fundus image quality assessment network according to any one of claims 1 to 6 to respectively obtain the global features and local features in the original retinal RGB image. Step S3, use the feature fusion module in the fundus image quality assessment network according to any one of claims 1 to 6 to fuse the obtained global features and local features, and obtain the final fused features. Step S4, input the final features into the fully connected layer and use them for the evaluation and classification tasks of the fundus image.

9. A fundus image quality assessment method combining global and local features according to claim 8, characterized in that, In step S2, when obtaining the enhanced local features by applying ClAHE to the G channel, apply ClAHE to the G channel of the original retinal RGB image.

10. A fundus image quality assessment method combining global and local features according to claim 8 or 9, characterized in that In step S3, provide independent supervision signals for the global feature extraction branch and the local feature extraction branch respectively through a three-way weighted loss strategy.

Citation Information

Cited By

  • Unmanned aerial vehicle image data enhancement method and system

    CN120725922A

  • Finger vein image quality evaluation method and system based on multi-feature interactive injection

    CN121095168A