Glaucoma Diagnosis Method Based on a Three-Stage Framework of Deep Sparse Autoencoder Ensemble
By adopting a three-stage framework integrated with deep sparse autoencoder in glaucoma diagnosis, the depth features and original features of fundus images are extracted and fused, and the problem of insufficient high-quality feature extraction in the prior art is solved, and the accuracy of glaucoma diagnosis is improved.
Patent Information
- Application Number
- CN202111626321.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The prior art is difficult to design appropriate features in glaucoma diagnosis, resulting in insufficient quality of feature extraction and difficulty in exploring the internal structural relationship between data.
Using a three-stage framework based on deep sparse autoencoder integration, the fundus image features are extracted through the DSA feature extraction model, and the original image features are fused with the depth features to obtain mixed features.
It improves the ability to dig out fundus images to hide feature structures, enhances the robustness and resolution of feature extraction, improves the discrimination of glaucoma, and improves the accuracy of diagnosis.
Smart Images

Figure CN114266763B_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a technology related to computer vision processing, and particularly relates to a glaucoma diagnosis method based on a three-stage framework integrated with a deep sparse autoencoder. Background Art
[0002] With the emergence of the Internet of Things, pioneering work has been carried out in patient health monitoring. Among them, automatic glaucoma diagnosis is one of them. Glaucoma is an eye disease caused by elevated intraocular pressure and improper drainage of the intraocular fluid. If left untreated, glaucoma will gradually and symptomlessly lead to loss of vision field and ultimately blindness. According to reports, the number of glaucoma patients globally will rise from 76 million in 2020 to 111.8 million in 2040. Since vision loss is irreversible, early detection and diagnosis will greatly prevent further degeneration of the optic nerve and have proven to reduce the blindness rate by about 50%.
[0003] In recent years, many automatic evaluation methods for the optic nerve head for glaucoma diagnosis have been proposed, which can be divided into two categories: based on clinical features and based on image features. Based on clinical features, a series of clinical evaluation indexes for glaucoma diagnosis, such as CDR, PPA, irregularity, etc., are designed to distinguish glaucoma images from normal images. Among these clinical features, CDR estimation is commonly used. Although CDR is the main index for diagnosing glaucoma, the quality change of fundus images and the presence of internal and surrounding blood vessels will affect the segmentation performance of OD and OC. To address this limitation, a series of traditional machine learning methods directly extract discriminative image features at the pixel level instead of extracting clinical features from the OD-OC segmentation region to explore the differences between normal images and glaucoma images. The extracted image features include wavelet decomposition, morphology-based, texture-based, and image descriptor-based, etc. After feature extraction, different classification strategies are selected for glaucoma diagnosis, such as artificial neural network (ANN), support vector machine (SVM), and k-nearest neighbor (KNN). However, since these features are handcrafted features, this process requires a large amount of expert knowledge, which is tedious. In addition, these feature extraction methods cannot find new high-quality features through multi-layer non-linear transformation and are difficult to explore the intrinsic structural relationship between data. Therefore, in traditional machine learning methods, how to design appropriate features is a difficulty and challenge.
[0004] In recent years, the Internet of Things has achieved a series of achievements in the medical industry. Among them, the glaucoma diagnosis method based on an end-to-end deep neural network has attracted much attention. However, these methods combine training features and classifiers, called co-adaptation. Therefore, the feature distribution should adapt to a specific decision boundary, thus learning a non-general data representation.
[0005] How to develop a glaucoma recognition method with reasonable feature extraction has become an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a glaucoma diagnosis method based on a three-stage framework integrated with deep sparse autoencoders to improve the accuracy of artificial intelligence in the glaucoma diagnosis process.
[0007] The technical solution provided by the present invention is specifically a glaucoma diagnosis method based on a three-stage framework integrated with deep sparse autoencoders, and the method includes the following steps:
[0008] S1: Obtain fundus images;
[0009] S2: Preprocess the fundus images to obtain preprocessed region-of-interest images and extract image features X, where X = [x 1 , x 2 ,..., x D , and D is the dimension of each region-of-interest image;
[0010] S3: Use the DSA feature extraction model to perform deep feature extraction on the image features X to obtain deep features d, where d = [d 1 , d 2 ,..., d z , z is the dimension of the features, and the expression of the DSA feature extraction model is:
[0011]
[0012] where is the reconstruction error, is the L2 weight regularization, is the sparse regularization, α is the L2 weight regularization coefficient, and β is the sparse regularization coefficient;
[0013] S4: Fuse the image features X and the deep features d to obtain hybrid features F, where F = {f 1 , f 2 ,..., f N}, and f i = [x i , d i ;
[0014] S5: After normalizing the hybrid features F, input them into the trained classification model for glaucoma diagnosis.
[0015] Preferably, in step S1, the fundus images are RGB images.
[0016] Further preferably, in step S2, the fundus image is preprocessed to obtain a preprocessed region of interest (ROI) image, specifically as follows:
[0017] S201: Input the fundus image into the optic disc rough segmentation model to locate the optic disc region of the fundus image and extract the optic disc center;
[0018] S202: Take the extracted optic disc center as the center of the region of interest, cut the fundus image to obtain the region of interest image;
[0019] S203: After correcting the region of interest image, extract the green channel image of the region of interest image to obtain the preprocessed region of interest image.
[0020] Further preferably, in step S201, the CHT algorithm is used to extract the optic disc center.
[0021] Further preferably, in step S203, the corrected RGB model and contrast-limited adaptive histogram equalization technology are used to correct the region of interest image.
[0022] Further preferably, in step S3, the value range of α is 0 to 1, and the value of β is 420.
[0023] Further preferably, in step S5, the formula for normalizing the mixed feature F is:
[0024]
[0025] where μ and σ are the mean and standard deviation vectors of the mixed feature respectively.
[0026] The glaucoma diagnosis method based on the three-stage framework integrated by the deep sparse autoencoder provided by the present invention extracts deep features from the original image features in the fundus image by using the DSA feature extraction model, and then fuses the original image features and the deep features to obtain the mixed feature, thereby improving the mining of the hidden feature structure of the fundus image, enhancing the robustness and resolution of the extracted features, and further improving the discrimination ability for glaucoma in the fundus image and the accuracy of diagnosis.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 Flowchart of a glaucoma diagnosis method based on a three-stage framework integrated with a deep sparse autoencoder provided by an open embodiment of the present invention;
[0031] Figure 2 Flowchart of step S2 in a glaucoma diagnosis method based on a three-stage framework integrated with a deep sparse autoencoder provided by an open embodiment of the present invention. Detailed implementation manners
[0032] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0033] To improve the accuracy of artificial intelligence in the process of glaucoma diagnosis, the present implementation provides a glaucoma diagnosis method based on a three-stage framework integrated with a deep sparse autoencoder. This method uses a non-end-to-end three-stage framework integrated with a deep sparse autoencoder to learn the generic data representation, improve the generalization ability of the model, enhance the complementarity between the deep features and the original features, and improve the resolution of the high-level feature expression.
[0034] See Figure 1 , the present implementation provides a glaucoma diagnosis method based on a three-stage framework integrated with a deep sparse autoencoder, which specifically includes the following steps:
[0035] S1: Obtain fundus images;
[0036] S2: Preprocess the fundus images to obtain the preprocessed region of interest (ROI) images and extract the image features X, where X = [x 1 , x 2 ,..., x D , and D is the dimension of each ROI image;
[0037] S3: Use the DSA feature extraction model to perform deep feature extraction on the image features X to obtain the deep features d, where d = [d 1 , d 2,..., d z , where z is the dimension of the feature, and the expression of the DSA feature extraction model is:
[0038]
[0039] where, is the reconstruction error, is the L2 weight regularization, is the sparse regularization, α is the L2 weight regularization coefficient, and β is the sparse regularization coefficient;
[0040] S4: Fuse the image feature X and the depth feature d to obtain a mixed feature F, where F = {f 1 , f 2 ,..., f N}, and f i = [x i , d i ;
[0041] S5: After normalizing the mixed feature F, input it into the trained classification model for glaucoma diagnosis.
[0042] Among them, in step S1, the fundus image is an RGB image.
[0043] In step S2: Preprocess the fundus image to obtain a preprocessed region of interest (ROI) image. See Figure 2 , specifically:
[0044] S201: Input the fundus image into the optic disc rough segmentation model to locate the optic disc region of the fundus image and extract the optic disc center. Among them, the above optic disc rough segmentation model is a trained existing optic disc rough segmentation model, and the CHT algorithm is used to extract the optic disc center when extracting the optic disc center;
[0045] S202: Use the extracted optic disc center as the center of the region of interest, cut the fundus image to obtain a region of interest image. Among them, the size of the region of interest image is the same, usually 170×170;
[0046] S203: After correcting the region of interest image, extract the green channel image of the region of interest image to obtain a preprocessed region of interest image. Among them, the correction of the region of interest image uses the modified RGB model and contrast-limited adaptive histogram equalization technology;
[0047] The main purpose of step S2 is to uniformly crop the region of interest image into the same size and extract the green channel in the RGB image to facilitate the depth feature extraction in step S3.
[0048] In step S3, the DSA feature extraction model is used to perform deep feature extraction on the image feature X to obtain a deep feature d, wherein the DSA feature extraction model introduces a mixed sparse feature group in SSAE to mine the hidden feature structure of the fundus image, thereby improving the robustness and resolution of the learned final high-level feature expression.
[0049] The specific DSA feature extraction model derivation process is as follows:
[0050] In the DSA representation extraction model, X∈R N×D Represents the input data obtained from the preprocessing stage, where N is the number of images and D is the dimension of each image. N×d represents hidden features, where is the output matrix of the tth hidden layer, d (t) is the number of units in the tth hidden layer. Assuming there are L AEs in the model, the number of hidden layers is also L. The hybrid feature can enhance the feature recognition ability by fusing the original feature and the encoding result of the previous AutoEncoder (AE) as the input of the next AE. The hybrid feature that can be formulated in Represents the connection operation that connects the original feature x with the network output feature h. E∈R (D+d)×d is a transformation matrix consisting of 0s and 1s. The elements of E are determined by the following factors:
[0051]
[0052] Where D is The covariance matrix, D mm is the mth diagonal element of D, v∈R d is a d-dimensional vector obtained by selecting the first d largest values among the diagonal elements of D, v n is the nth element of v.
[0053] Since the divergence of features can be used to measure the classification resolution ability, the objective function of the mixed features can be defined as:
[0054]
[0055] The encoder function of the pth AE in DSA can be defined as:
[0056]
[0057] In the formula, is the (p-1)th mixed feature (the original feature x is mixed with h (p-1) The reconstruction result output by connecting the encoding result of W p1 and bp1 are the weight matrix and bias vector of the encoder in the p-th AE. f(·) is the logistic regression sigmoid function.
[0058] Similarly, the decoder function of the p-th AE is:
[0059]
[0060] where is the reconstruction result of the (p - 1)-th mixed feature output. W p2 and b p2 represent the weight matrix and bias vector of the decoder in the p-th AE.
[0061] To explore the important structures for glaucoma diagnosis, a sparse constraint is introduced during the AE training process. Similar to the original SAE, KL divergence is used for sparse representation in this implementation, specifically expressed as:
[0062]
[0063] where ρ is the sparse parameter, and ρ K is the average activation value of all training samples on the K-th hidden neuron.
[0064] In addition, since the neurons in the hidden layer are interdependent, it is necessary to explore the relationships between them. In this implementation, local dependence is adopted, and the hidden feature H is divided into two sub-parts, including the original feature H G1 and the output feature H G2 of the previous layer, that is, H = [H G1 , H G2 . Then the sparse constraint on the hidden feature H can be expressed as:
[0065] Γ(H) = ||H G1 || 1 + ||H G2 || 1 (6)
[0066] where ||·|| 1 is the L1 norm. By introducing the sparsity of the hidden feature representation, the ability of the hidden feature representation is enhanced, and at the same time, the dependence relationships between different parts are suppressed, making them compete with each other. Combining the sparse constraint with the mixed features into a unified framework, a mixed sparse feature group can be obtained. The final objective function of the p-th AE in DSA is:
[0067]
[0068] In the formula, it contains three terms. The first term is the reconstruction error, the second term is the L2 weight regularization, and the third term is the sparse regularization. α and β are the two coefficients of the L2 weight regularization and the sparse regularization respectively.
[0069] The deep features are obtained through the above DSA feature extraction model, denoted as d = [d 1 , d 2 ,..., d z .
[0070] In step S4, the original image features X and the deep features d are fused to obtain the mixed features F, where F = {f 1 , f 2 ,..., f N}, and f i = [x i , d i , where each original image X = [x 1 , x 2 ,..., x D ∈ R 1×D , and the mixed feature representation vector f i = [x i , d i ∈ R 1×(D+Z) .
[0071] In step S5, after normalizing the above mixed features F, they are input into the trained classification model for glaucoma diagnosis. The formula for normalizing the mixed features F is: In the formula, μ and σ are the mean and standard deviation vectors of the mixed features respectively. The classification model can adopt an existing classification model.
[0072] The glaucoma diagnosis method provided by the above implementation scheme extracts deep features from the original image features in the fundus image through the DSA feature extraction model, and then fuses the original image features and the deep features to obtain the mixed features, thereby improving the mining of the hidden feature structure of the fundus image, enhancing the robustness and resolution of the extracted features, and further improving the discrimination ability for glaucoma in the fundus image and the accuracy of diagnosis.
[0073] In the above DSA feature extraction model, the L2 weight regularization coefficient α and the sparse regularization coefficient β can be trained on the fundus image samples in the database to obtain the optimal parameters to reduce the reconstruction error, thereby obtaining the expression features and improving the representation ability of the class. Therefore, the softmax layer is regarded as the classification layer connected to the top of the DSA, and the entire network is fine-tuned in a supervised manner.
[0074] For each original image X = [x 1 , x2 ,..., x D ∈ ℝ 1×D , after training the entire network, each hidden layer has a new feature vector. As is well known, shallower layers have simple features, and higher layers have complex features inherent in the data itself, which is beneficial for classification. Therefore, this implementation selects the output of the last hidden layer as the deep feature, denoted as d = [d 1 , d 2 ,..., d z . Finally, taking the original feature X = [x 1 , x 2 ,..., x D , and the obtained deep feature d = [d 1 , d 2 ,..., d z , considering constructing a hybrid feature representation vector f i = [x i , d i ∈ ℝ 1×(D+Z) , the hybrid features of all enhanced region of interest images obtained can be expressed as:
[0075] F = {f 1 , f 2 ,..., f N} (8)
[0076] where f i ∈ ℝ 1×(D+Z) represents the feature of the i-th region of interest image, and N is the total number of region of interest images.
[0077] To verify the effectiveness of the method provided by the above implementation, the following experiments were conducted, specifically as follows:
[0078] In the experiment, we used two publicly available databases (REFUGE and Drishti - GS1) to verify the effectiveness of the proposed method. Table 1 shows the brief information of the used databases.
[0079] REFUGE database: It stores 1200 color retinal fundus images in JPEG format, among which, there are 120 glaucoma images and 1080 normal images. There are three subsets in the REFUGE database, which contain 400 training images, 400 validation images, and 400 test images respectively. Since only 400 training images and 400 validation images can be downloaded from the Internet, in this experiment, we used these available images to train and test the proposed method.
[0080] Drishti - GS1 Database: It contains 101 uncompressed retinal fundus images with a 30 - degree field of view. Among them, 31 are normal field of view images and 70 are glaucoma field of view images. For each fundus image, the corresponding ground truth is marked by 4 glaucoma ophthalmologists. Since the database only contains 50 training images, when the network training data is too little, the network may lead to overfitting. To solve this problem, this solution uses a data augmentation method based on translation and rotation to expand the training images to prevent this problem from occurring. Among them, the images in the Drishti - GS1 database are expanded to 800, with 600 for training and 200 for testing.
[0081] Table 1 shows the statistical data of the database used to evaluate the proposed method.
[0082]
[0083] In the experiment, three widely used evaluation metrics, including accuracy, sensitivity, and specificity, are adopted to verify the effectiveness of the proposed method. Their mathematical expressions are as follows:
[0084]
[0085]
[0086]
[0087] True positives (tp) and true negatives (tn) represent the number of images correctly identified as glaucoma and non - glaucoma; false positives (fp) and false negatives (fn) are the number of incorrectly detected images among glaucoma and non - glaucoma images.
[0088] In addition, the receiver operating characteristic curve (ROC) is also used to compare the diagnostic performance of glaucoma. The vertical axis of the ROC curve is sensitivity, and the horizontal axis is (1 - sensitivity). The area under the ROC curve is denoted by AUC, and AUC can be used to measure and describe the performance of different methods. The higher the AUC, the better the performance of the method.
[0089] All experiments were conducted in a unified experimental environment with Windows 10 operating system, Intel(R) i7-7700 CPU @ 2.8GHz, 16GB random access memory (RAM), and NVIDIA 1050Ti GPU. The programming tool was MATLAB, version R2020b. To obtain more useful information from the limited database and reduce overfitting, five-fold cross-validation was adopted in this scheme. For the REFUGE and Drishti-GS1 databases, we first resized the ROI images to 64×64 pixels to save computational complexity. Next, two-thirds of all the images were randomly selected as training data, and the rest were used as test data. Therefore, 600 images were randomly selected for training, and the remaining 200 images were used for testing. The sample selection process was repeated five times, and the average accuracy was regarded as the final performance.
[0090] In this experiment, we set the number of AEs for the REFUGE and Drishti-GS1 databases to 3, and the sparse regularization β value to 420. Since the number of each hidden unit is related to the original data, this scheme adopted the method of grid search to find the optimal structure. We searched the grids {60, 80, 100, 120}, {30, 40, 50, 60}, and {10, 14, 18, 22, 26, 30} for the first AE (AE1), the second AE (AE2), and the third AE (AE3) respectively to optimize the number of hidden units. The best ones were used in the following experiments. It can be seen from the experimental results that when the number of hidden units of AEs {AE1, AE2, AE3} is {100, 40, 30} and {60, 60, 18} respectively, the proposed method can achieve the best classification performance (ROC-AUC) on the REFUGE database (AUC = 96.94%) and the Drishti-GS1 database (AUC = 98.28%).
[0091] There are two main parameters in the method proposed in the above implementation scheme, namely the sparsity ratio ρ and the L2 weight regularization α. Therefore, how to properly set their values is crucial for this method. In the experiment, for the REFUGE database and the Drishti-GS1 database, we adjusted the value of the ρ parameter by searching the grid {10-5, 10-4, 10-3, 10-2, 10-1, 1, 10, 100}, and the value range of the α parameter was 0 to 1 with an interval of 0.05.
[0092] First, we tested the classification performance (ROC-AUC) of the proposed method at different ρ values on the REFUGE and Drishti-GS1 databases. When the sparsity ratio of all AEs was set to 0.85, the proposed glaucoma diagnosis method achieved the maximum AUC value on the REFUGE and Drishti-GS1 databases.
[0093] Then, we examined the classification performance of the method at different parameter α values. As the value of parameter α increased, the classification performance of the proposed method became better, indicating that weight decay plays an important role in deep feature extraction. However, after reaching the optimal performance, the classification performance decreased as the value of α increased. The main reason is that excessive regular penalties will cause too many connection weights to be zero. Therefore, in this scheme, the L2 weight regularization parameters for the REFUGE database and the Drishti-GS1 database were set to 0.1 and 1, respectively.
[0094] Finally, we will analyze the performance of the proposed method during the training process on the REFUGE and Drishti-GS1 databases. The proposed method can obtain a very small mean squared error on both databases. The results show that the method has generalized the information in the training images. In other words, the model has been well trained and can be tested. We set the MaxEpochs for the REFUGE database and the Drishti-GS1 database to 100 and 65, respectively.
[0095] In summary, Table 2 shows the final parameter settings of the method proposed in our experiment.
[0096] Table 2 Parameter settings of the method proposed in the experiment.
[0097]
[0098] We will test and analyze the performance of different features under different classifiers. Five commonly used classifiers, namely SVM, ELM, KNN, RF, and Softmax, were selected for the experiment. Among them, Softmax is a supervised classifier, which will be combined in the later stage to distinguish the output of the SSAE into various categories, referred to as end-to-end. The features used are original features, deep features, and hybrid features. The original features refer to the output images obtained in the preprocessing stage. The deep features are the output of the original SSAE. The hybrid features are the combination of the output of the proposed DSA method and the original features.
[0099] Table 3 Classification accuracies (%) of different classifiers under different features in the REFUGE database
[0100]
[0101] Table 4 Classification accuracy (%) of different classifiers under different features in the Drishti-GS1 database
[0102]
[0103] Tables 3 and 4 show the classification accuracies of different classifiers under different features. From Tables 3 and 4, we can learn the following two points. On the one hand, the hybrid features are more robust and effective than other features under all classifiers and can achieve better classification performance. For the REFUGE database, the accuracy of the SVM hybrid features is 2% higher than that of the deep features and 3% higher than that of the original features. Similarly, the hybrid features can also achieve the best classification performance superior to other features on the Drishti-GS1 database. On the other hand, stable classification results can be obtained by using the hybrid features of different classifiers. This phenomenon not only indicates that the hybrid features are more robust and discriminative than other features, but also that the proposed non-end-to-end method hinders the adaptation of the feature extractor to a specific strong classifier, improving the generalization ability of the model. Under the hybrid features, the classification accuracies of SVM and RF are both the highest. When the number of decision trees is large, the training of RF requires more space and time, so the efficient support vector machine is a better choice to verify the effectiveness of the proposed method. Through the above analysis, we can conclude that the use of hybrid features and support vector machine classifiers is beneficial to improving the diagnostic performance of glaucoma. Therefore, we will use this combination for subsequent performance comparison.
[0104] To verify the effectiveness of the method provided by the above implementation, we compare the method of this implementation with the state-of-the-art glaucoma diagnosis methods, including traditional machine learning-based methods (AWLCSC, wavelet features, superpixel segmentation, semi-supervised clustering, feature fusion) and deep learning-based methods (SAE, SSAE, SSAE+Softmax, GoogleNet, SSAE+LASSO, CNNs integration, Inception v3). Tables 5 and 6 respectively describe the best classification accuracies and AUCs of different methods on the REFUGE and Drishti-GS1 databases.
[0105] Table 5 Comparison of the method of this implementation with the latest methods on the REFUGE database
[0106]
[0107] Table 6 Comparison of the method of this implementation with the state-of-the-art methods on the Drishti-GS1 database
[0108]
[0109] From the above results shown in Table 5 and Table 6, the following points can be seen:
[0110] (1) Using traditional machine learning methods (image feature - based or CDR - based) to classify ROI images for glaucoma diagnosis. The classification accuracies and AUCs on the REFUGE database are (58.70%, 57.20%), (63.53%, 64.23%), (81.95%, 82.00%), (85.54%, 84.00%) and (77.67%, 78.02%), (88.63%, 87.00%) respectively. Since the features extracted in these studies are shallow features and rely on empirical knowledge, they cannot effectively mine the inherent non - linear complex relationships between high - dimensional fundus images. Therefore, the obtained results are not very promising.
[0111] (2) Applying the CNN model to GoogleNet, Ensemble of CNNs, and Inception v3 achieved good performance. The main reason is that the features learned through multi - layer neural networks provide more important data descriptions, which helps glaucoma diagnosis. However, as is well known, the CNN model has more hyperparameters to be set. As the number of layers increases, the number of hyperparameters in the CNN model will greatly increase the complexity of the model. In addition, the CNN model can achieve excellent performance on large databases. Therefore, when the number of samples is small, its generalization ability will be correspondingly reduced.
[0112] (3) Different from CNN, SAE, SSAE, SSAE + LASSO, and SSAE + Softmax combine unsupervised layer - by - layer pre - training and supervised fine - tuning to learn expression features and can obtain better performance than CNN. Although SAEs can explore the non - linear structure of data, these methods only use the original auto - encoder structure to improve performance while ignoring the encoder structure, weakening the feature - learning ability. In addition, when the number of samples is small, SAEs are prone to overfitting, which limits their generalization ability to learn effective features. Moreover, through the joint optimization of the feature extractor and the classifier, both SAEs and CNN are end - to - end models, but this process results in a feature distribution that is too complex for a specific classifier, reducing the test performance.
[0113] (4) Among all the comparison methods, the proposed solution of this application has the best performance on the REFUGE and DRISHTI-GS1 databases. The main reasons are as follows: 1) Different from the existing end-to-end deep learning methods, the proposed solution of this application presents a three-stage framework of deep sparse autoencoder integration for glaucoma diagnosis. This method relaxes the strict limitation that complex feature distributions must adapt to specific classifiers, can learn general data representations, and improves the generalization ability of the model. 2) The original features are fused into the hidden layer to construct hybrid features, enhancing the complementarity between the original features and the deep features. Therefore, as the number of hidden layers increases, the original information will be retained in the high-level expressions, improving the robustness and recognition rate of the learned features. 3) A sparse constraint is introduced into the hybrid features to construct a group of hybrid sparse features, further reducing the influence of noise and redundant features. In summary, the three-stage framework of deep sparse autoencoder integration for glaucoma diagnosis proposed by the solution of this application is helpful for glaucoma diagnosis.
[0114] Other embodiments of the present invention will be readily apparent to those skilled in the art upon considering the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention following the general principles of the invention and including known common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only illustrative, and the true scope and spirit of the invention are pointed out by the following claims.
[0115] It should be understood that the present invention is not limited to what has been described above and can be modified and changed without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders, characterized in that, it includes the following steps: S1: Obtain fundus images; S2: Preprocess the fundus image to obtain a preprocessed region-of-interest image and extract image features X, where X = [x 1 , x 2 ,..., x D , and D is the dimension of each region-of-interest image; S3: Use the DSA feature extraction model to perform deep feature extraction on the image feature X to obtain the deep feature d. The derivation process of the DSA feature extraction model is as follows: In the DSA representation extraction model, X ∈ R N×D represents the input data obtained from the output of the preprocessing stage, where N is the number of images and D is the dimension of each image; H ∈ R N×d represents the hidden feature, where is the output matrix of the t-th hidden layer, and d (t) is the number of units in the t-th hidden layer; assuming there are L AEs in the model, then the number of hidden layers is also L; the hybrid feature is formulated by fusing the original feature and the encoding result of the previous autoencoder as the input of the next AE where represents the concatenation operation that concatenates the original feature x and the network output feature h; E ∈ R (D+d)×d is a transformation matrix composed of 0 and 1, and the elements of E are determined by the following factors: where D is the covariance matrix, D mm is the m-th diagonal element of D, v ∈ R d is a d-dimensional vector obtained by selecting the top d maximum values among the diagonal elements of D, and v n is the n-th element of v; Use the divergence of features to measure the classification resolution ability, and the objective function of the mixed features is defined as: The encoder function of the p-th AE in DSA is defined as: In the formula, is the reconstruction result of the (p - 1)-th hybrid feature output, where the original feature x is concatenated with the encoded result of h (p-1) , W p1 and b p1 are the weight matrix and bias vector of the encoder in the p-th AE; f(·) is the logistic regression sigmoid function; Similarly, the decoder function of the p-th AE is: In the formula, is the reconstruction result of the (p - 1)-th hybrid feature output; W p2 and b p2 represent the weight matrix and bias vector of the decoder in the p-th AE; Use KL divergence for sparse representation, specifically expressed as: where ρ is the sparsity parameter, and ρ K is the average activation value of all training samples on the K-th hidden neuron; Using local dependence, the hidden feature H is divided into two sub-parts, including the original feature H G1 and the output feature H of the previous layer G2 , that is, H = [H G1 , H G2 . Then the sparse constraint on the hidden feature H is expressed as: Γ(H) = ||H G1 || 1 + ||H G2 || 1 (6) where, ||·|| 1 is the L1 norm; combining the sparse constraint and the mixed features into a unified framework to obtain a mixed sparse feature group; the final objective function of the p-th AE in DSA is as follows: In the formula, there are three terms. The first term is the reconstruction error, the second term is the L2 weight regularization, and the third term is the sparsity regularization. α and β are the two coefficients of the L2 weight regularization and the sparsity regularization respectively. Among them, d = [d 1 , d 2 ,..., d z , z is the dimension of the feature, is the reconstruction error, is the L2 weight regularization, is the sparsity regularization, α is the L2 weight regularization coefficient, and β is the sparsity regularization coefficient; The deep features are obtained by the above DSA feature extraction model, denoted as d = [d 1 , d 2 ,..., d z ; S4: Fuse the image feature X with the depth feature d to obtain a hybrid feature F, where F = {f 1 , f 2 ,..., f N}, and f i = [x i , d i ; S5: After normalizing the mixed feature F, input it into the trained classification model for glaucoma diagnosis.
2. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 1, characterized in that, in step S1, the fundus image is an RGB image.
3. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 1, characterized in that, in step S2, preprocess the fundus image to obtain a preprocessed region of interest image, specifically: S201: Input the fundus image into the optic disc rough segmentation model to locate the optic disc region of the fundus image and extract the optic disc center; S202: Use the extracted optic disc center as the center of the region of interest, cut the fundus image to obtain the region of interest image; S203: After correcting the region of interest image, extract the green channel image of the region of interest image to obtain the preprocessed region of interest image.
4. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 3, characterized in that, in step S201, the CHT algorithm is used to extract the optic disc center.
5. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 3, characterized in that, in step S203, the corrected RGB model and contrast-limited adaptive histogram equalization technology are used to correct the region of interest image.
6. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 1, characterized in that, in step S3, the value range of α is 0 to 1, and the value of β is 420.
7. The glaucoma diagnosis method based on a three-stage framework integrating deep sparse autoencoders according to claim 1, characterized in that, in step S5, the formula for normalizing the mixed feature F is: where μ and σ are the mean and standard deviation vectors of the mixed features respectively.