Ophthalmology image classification method based on selective state space fusion
By introducing the S3FNet method of selective state space fusion in ophthalmic image analysis, combining wavelet multi-scale feature extraction, fusion of CNN and Transformer, and visual-language cross-modal learning, the problem of the existing technology processing complex data and lack of cross-modal information fusion is solved, and efficient ophthalmic disease classification and analysis is achieved.
Patent Information
- Application Number
- CN202510408322.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-16
AI Technical Summary
Existing ophthalmic image analysis techniques perform poorly in processing high-dimensional complex data, capturing multi-scale information, and coping with noise and redundant information in images, and lack effective cross-modal information fusion.
A ophthalmic image classification method based on selective state space fusion is proposed, called S3FNet. This method overcomes many technical bottlenecks in traditional methods by combining wavelet multi-scale feature extraction, effective fusion of CNN and Transformer, and visual-language cross-modal learning.
It significantly improves the efficiency and accuracy of fundus medical image feature extraction, improves the classification accuracy and robustness of ophthalmic diseases, can fully tap into the multi-scale frequency domain features in fundus images, and captures global information through the Transformer encoder.
Smart Images

Figure CN120014691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision deep learning and medical image analysis, and specifically to an ophthalmic image classification method based on selective state space fusion. Background Art
[0002] With the rapid development of medical imaging technology, the early detection of ophthalmic diseases has become an important field in modern medical research and clinical treatment. In particular, fundus medical images (such as retinal images) play a vital role in the discovery of ophthalmic diseases such as diabetic retinopathy (DR), glaucoma, cataracts, and macular degeneration (AMD). However, existing fundus medical image analysis methods still have some limitations in disease identification, especially when processing high-dimensional complex data, capturing multi-scale information, and dealing with noise and redundant information in images. The performance of traditional methods is not ideal.
[0003] Traditional ophthalmic disease image classification usually relies on manual feature extraction and regularized machine learning algorithms, which are difficult to perform well in complex and diverse fundus images. In addition, although convolutional neural networks (CNNs) have advantages in local feature extraction, they have certain limitations in capturing long-distance dependencies and global contextual information. To overcome this problem, in recent years, Transformer-based models have gradually emerged in computer vision tasks, especially in global modeling and capturing long-distance dependencies. However, Transformer also faces problems of computational efficiency and feature fusion.
[0004] In addition, most current ophthalmic image analysis systems focus on single processing of image data and lack effective cross-modal information fusion. Ophthalmic disease image classification not only depends on the image itself, but may also be affected by the patient's health background, clinical records and other multimodal information. Summary of the invention
[0005] In order to solve the shortcomings of existing ophthalmic image analysis technology, the present invention proposes an ophthalmic image classification method based on selective state space fusion, the selective state space fusion vision-language network (S3FNet). The network overcomes many technical bottlenecks in traditional methods by combining wavelet multi-scale feature extraction, effective fusion of CNN and Transformer, and visual-language cross-modal learning. In particular, the introduction of the Mamba architecture based on the selective state space model (SSM) further improves the efficiency and accuracy of fundus medical image feature extraction, and provides a new solution for the automated and intelligent analysis of ophthalmic diseases. Through this method, efficient analysis of fundus medical images is achieved with high classification accuracy and robustness.
[0006] To achieve the above object, the technical solution of the present invention is:
[0007] An ophthalmic image classification method based on selective state space fusion comprises the following steps:
[0008] Step 1: Obtain fundus medical image data, collect fundus medical images from different patients, perform preprocessing, and generate a fundus medical image data set. Specifically: fundus medical image data collection and annotation; standardization processing to make the color value of the image conform to a unified range; adaptive histogram equalization to enhance the details in the image; rotation, scaling, flipping, and cropping to expand the data set; finally, divide the data set into training set, validation set, and test set.
[0009] Step 2: Build the S3FNet model. This model achieves efficient analysis of fundus medical images through the collaborative work of three main modules: the wavelet multi-scale feature extractor first extracts multi-scale features from the image, which are then used by the inductive bias transformer encoder (IBTE) to capture long-distance dependencies, and finally the multi-model aggregation module (MA) combines vision-language joint learning to improve classification performance. Specifically, it includes:
[0010] (1) The wavelet multiscale feature extractor is used to extract multiscale features of different frequencies from fundus medical images. The implementation steps are as follows:
[0011] The input fundus medical image is subjected to wavelet transform, and the different frequency information in the image is extracted by multi-level decomposition, and the feature maps of multiple scales are obtained through the convolutional network CNN. The high-frequency part extracted by wavelet transform captures the edge, texture and other details of the image; the low-frequency part extracts the background information and structural features.
[0012] Finally, the formed multi-scale feature map representation is input into the Mamba architecture based on the selective state space model (SSM). The efficient feature extraction capability of the Mamba architecture is utilized to further optimize the feature representation and enhance the feature discrimination and expression ability.
[0013] The features processed by the Mamba architecture are spliced with the input fundus medical image before processing to generate a multi-level feature map, thereby realizing the extraction of multi-scale features of different frequencies.
[0014] (2) The Inductive Biased Transformer Encoder (IBTE) fuses local features and global features through a cross-attention mechanism to capture the long-range dependencies of fundus images. The implementation steps are as follows:
[0015] The input fundus medical image is decomposed into multiple small blocks by patching, and each small block is converted into a feature vector.
[0016] For each feature vector, the spatial position information is enhanced using the position encoding information to form a query sequence.
[0017] These query sequences are cross-attentioned with the multi-level feature maps (Key and Value) output by the wavelet multi-scale extractor to obtain global features. The feature maps of multiple scales extracted by CNN and the global features learned by Transformer are spliced and combined.
[0018] The global context modeling of the input fundus medical image is performed through a multi-layer Transformer module to further extract the global visual features of the image.
[0019] (3) The multi-model aggregation module (MA) combines vision-language joint learning to improve the classification performance of ophthalmic diseases. The implementation steps are as follows:
[0020] Using the contrastive language-image pre-training model CLIP, the visual features of fundus medical images are fused with the corresponding text descriptions (such as "normal", "diabetes", "glaucoma", etc.).
[0021] Through contrastive learning, the similarity between visual features and corresponding text descriptions is calculated, enabling the model to understand the relationship between images and text.
[0022] For each category (such as diabetic retinopathy, cataract, etc.), a corresponding text description is constructed as prompt information for visual-language learning.
[0023] The visual embedding and text embedding are matched by cosine similarity calculation, and the final classification probability is generated by the Softmax normalization method.
[0024] Step 3: Train the S3FNet model using the training set in the fundus medical image dataset and optimize the network parameters to achieve the best classification effect. The specific steps are as follows:
[0025] Training set construction: The processed training set is input into the S3FNet model for training and optimization of model parameters.
[0026] Optimization strategy: Adam optimizer is used for parameter update, the initial learning rate is set to 0.001, and the learning rate decay strategy is used to reduce the learning rate every 20 epochs of training.
[0027] Batch size: Choose an appropriate batch size (such as 32) to stabilize the training process and accelerate gradient descent.
[0028] Regularization: Prevent overfitting through L2 regularization (weight decay).
[0029] Early stopping strategy: Use the early stopping strategy to stop training early when the performance of the validation set does not improve to avoid overfitting.
[0030] Parameter tuning: Optimize hyperparameters (such as learning rate, batch size, etc.) through cross-validation and grid search to ensure that the model achieves optimal performance.
[0031] Step 4: Use the test set in the fundus medical image dataset to evaluate the trained S3FNet model, evaluate the model through the test set, and generate the final classification model. The specific steps are as follows:
[0032] Model evaluation: Use the test set to evaluate the trained S3FNet model, and comprehensively evaluate the model performance through indicators such as accuracy, F1 score, and AUC.
[0033] Classification results: Output the final ophthalmic disease image classification results. The classification types include but are not limited to diabetic retinopathy (DR), glaucoma, cataract, AMD, hypertension, etc.
[0034] Step 5: Use the trained S3FNet model to classify ophthalmic diseases. Use the final trained S3FNet model to classify fundus medical images and output the type of ophthalmic diseases.
[0035] Input test image: Input the fundus medical image to be classified into the trained S3FNet model.
[0036] Feature extraction and classification: The frequency domain and spatial features of the image are extracted through the wavelet multi-scale feature extractor. Global feature modeling is performed through the inductive bias Transformer encoder. The multi-model aggregation module is used for visual-language joint learning to generate the final classification result.
[0037] Compared with the existing invention, the beneficial effects of the present invention are:
[0038] The method based on selective state space fusion vision-language network (S3FNet) proposed in the present invention significantly improves the efficiency and accuracy of fundus medical image feature extraction by combining wavelet multi-scale feature extraction, effective fusion of convolutional neural network (CNN) and Transformer, and visual-language cross-modal learning. In particular, the Mamba architecture based on the selective state space model (SSM) is introduced to intelligently select key information and compress redundant features, thereby improving computational efficiency and reducing processing time. The present invention can not only fully mine the multi-scale frequency domain features in fundus images, but also capture global information through the Transformer encoder, and combine the CLIP model for visual-language fusion learning, which significantly improves the classification accuracy of ophthalmic diseases. In addition, the entire analysis process effectively reduces manual intervention through automated training, optimization and evaluation. Therefore, the present invention shows the broad application prospects of deep learning and multimodal fusion technology in the field of medical image analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the overall structure of the present invention;
[0040] Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0041] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0042] The present invention provides an ophthalmic image classification method based on selective state space fusion, such as Figure 1 and Figure 2 As shown, the specific operations are as follows:
[0043] Obtain fundus medical image data, collect fundus medical images from different patients, perform preprocessing, and generate a fundus medical image dataset. The specific steps are as follows:
[0044] Data collection: Collect fundus medical imaging data from multiple hospitals, covering a variety of ophthalmic diseases, such as diabetic retinopathy, glaucoma, cataracts, macular degeneration, etc., as well as normal fundus images.
[0045] Data annotation: Professional ophthalmologists annotate each fundus medical image in the dataset, including information such as disease category and disease severity.
[0046] Color normalization: Through standardization, the color values of the image are made to conform to a unified range, reducing the color differences between different devices.
[0047] Color-adaptive histogram equalization (CLAHE): Enhances details in the image, especially low-contrast areas, and improves the recognition of image features.
[0048] Data enhancement: Expand the data set through rotation, scaling, flipping, cropping and other techniques to improve the robustness and generalization ability of the model.
[0049] Dataset division: The dataset is divided into training set (70%), validation set (15%) and test set (15%) to ensure the stability of training and evaluation.
[0050] 1. Implementation of Wavelet Multiscale Feature Extractor
[0051] In the present invention, wavelet transform is used as a pre-processing module for feature extraction. Different frequency information in fundus medical images is extracted through multi-level wavelet transform to form a multi-scale feature map. The specific implementation steps are as follows:
[0052] Wavelet transform decomposition: Perform multi-level wavelet transform on the input fundus medical image to decompose the image into low-frequency and high-frequency parts. The low-frequency part is mainly used to capture the overall structural information of the image, while the high-frequency part is used to capture the detailed features in the image (such as edges and textures). Select the appropriate wavelet basis function and set the number of wavelet decomposition layers according to the specific characteristics of the fundus image to extract the required frequency information.
[0053] Local feature extraction: The local features in the multi-scale feature map generated after wavelet transform are extracted through convolutional neural network (CNN). The details in the image are captured by processing feature maps of different scales.
[0054] Multi-scale feature fusion: Feature maps of different scales are spliced to generate feature representations containing multi-scale information, providing rich local feature input for subsequent global modeling.
[0055] 2. Implementation of Inductive Biased Transformer Encoder (IBTE)
[0056] In order to effectively capture the global dependencies in fundus medical images and improve the performance of the model, the present invention adopts the Inductive Bias Transformer Encoder (IBTE). It mainly fuses local features with global features through the cross-attention mechanism. The specific implementation steps are as follows:
[0057] Image block segmentation: The input fundus medical image is divided into multiple small patches (patches). Each small patch will be converted into a feature vector for subsequent processing. A position code is added to each small patch to retain its spatial position information to form a query sequence (Query).
[0058] Cross-attention mechanism: The local features (Key and Value) obtained from the wavelet multi-scale feature extractor are cross-attended with the query sequence (Query) generated by the image block. Through cross-attention, the model can capture the relationship between local and global features in a global scope.
[0059] Global information fusion: After being processed by multiple layers of Transformer encoders, the global visual features of fundus medical images are gradually learned and extracted. The output of each layer is normalized and optimized, allowing the model to extract global dependency information more accurately.
[0060] 3. Implementation of Multi-Model Aggregation Module (MA)
[0061] In the multi-model aggregation module (MA), the present invention combines the CLIP model for visual-language joint learning. Through the contrastive learning mechanism, the CLIP model can combine the visual features of fundus medical images with text descriptions related to ophthalmic diseases, thereby improving classification accuracy. The specific implementation steps are as follows:
[0062] Vision-language fusion: Using the pre-trained CLIP model, the visual features of fundus medical images and the corresponding text descriptions (such as "diabetic retinopathy", "glaucoma", "cataract", etc.) are compared and learned. Through comparative learning, the similarity between the image and the text is maximized, so that the model can understand the relationship between visual information and language information.
[0063] Text prompt construction: Build a text description template for each ophthalmic disease (such as diabetic retinopathy, macular degeneration, etc.) as a prompt for visual-language joint learning. This step can further guide the model to accurately identify the disease category.
[0064] Classification output: By calculating the cosine similarity of image embedding and text embedding and normalizing through Softmax, the classification probability of each ophthalmic disease category is finally generated.
[0065] The wavelet multiscale feature extractor uses wavelet transform for image decomposition. Wavelet transform is an effective frequency domain analysis tool that can extract high-frequency and low-frequency information in fundus medical images while retaining spatial details. Through multi-level wavelet transform, we can extract multi-scale features from the image, capture details such as the edge and texture of the image, and then help the model identify the subtle features of ophthalmic diseases. Especially in fundus medical images, subtle visual differences often contain information that helps with classification, so the use of wavelet transform can improve the accuracy of feature extraction and further improve the classification performance of the model.
[0066] The Inductive Biased Transformer Encoder (IBTE) fuses local and global features through a cross-attention mechanism. Traditional convolutional neural networks (CNNs) perform well in capturing local features, but have limitations in modeling global contextual relationships. To this end, IBTE introduces a cross-attention mechanism to fuse the local features extracted by CNN with the global features learned in the Transformer, thereby effectively making up for the shortcomings of capturing local and global features. In this way, the model can more accurately understand the complex structures and long-distance dependencies in fundus medical images. The cross-attention mechanism better expresses the relationship between local and global features through fine weight allocation.
[0067] The multi-model aggregation module (MA) combines CLIP to perform visual-language joint learning. In order to further improve the classification ability of ophthalmic diseases, the present invention adopts the CLIP model to fuse the visual features of fundus medical images with disease-related text descriptions through contrast learning. The CLIP model has a strong cross-modal learning ability and can effectively associate images and texts, thereby enhancing the model's comprehensive understanding of ophthalmic image classification. Combining visual and language information can not only improve the accuracy of classification, but also help the model better cope with different categories of disease manifestations. Therefore, visual-language joint learning provides S3FNet with richer semantic information, further improving the classification performance.
[0068] The selective state space model (SSM) and Mamba architecture. The Mamba architecture is based on the selective state space model (SSM) and improves the efficiency of feature extraction by introducing a selective scanning mechanism. Unlike traditional global processing methods, the selective scanning mechanism can intelligently select key information and compress redundant features, thereby improving computational efficiency and reducing processing time. The Mamba architecture is particularly suitable for processing the complexity of high-dimensional data in fundus medical images, and can process large amounts of image data quickly and efficiently while maintaining high precision. Through a parallel scanning algorithm, Mamba is able to process long sequence data, avoid computational bottlenecks that may occur in traditional methods, and improve the operating efficiency and performance of the entire model.
[0069] The dataset preprocessing method includes color normalization, histogram equalization (CLAHE) and data enhancement. In fundus medical image analysis, due to imaging differences between different devices, there may be significant differences in the brightness, contrast and color distribution of the image, which affects the training effect of the model. In order to eliminate these effects, color normalization and adaptive histogram equalization (CLAHE) are used in the preprocessing process. These methods can unify the color space of the image, improve the contrast of the image, and enhance the details in the image. In addition, data enhancement techniques such as rotation, scaling, flipping, etc. can expand the dataset and help the model improve its tolerance to image deformation and noise, thereby improving the robustness and generalization ability of the model.
[0070] The training and optimization strategies include learning rate scheduling, momentum optimization, batch size adjustment, etc. In order to ensure the efficiency and stability of the training process, the present invention adopts a number of optimization strategies. Through the learning rate scheduling strategy, the learning rate is gradually attenuated to avoid overfitting and improve the convergence speed of the model. The momentum optimization method accelerates the process of gradient descent, avoids possible shocks during training, and improves training efficiency. The adjustment of batch size ensures the stability of each update and avoids the instability caused by too small batches. In addition, the early stopping strategy can stop training when the performance of the verification set is not improved, avoiding overfitting while saving computing resources. The combination of all these strategies ensures that the S3FNet model is both efficient and stable during the training process, and can achieve the best classification effect.
[0071] The process uses efficient computing and hardware acceleration. In order to speed up the model training, the present invention is trained in a high-performance computing environment, using a server equipped with an Nvidia GeForce GTX 1080 GPU. GPU acceleration can not only significantly increase the training speed, but also reduce computing bottlenecks when processing large-scale data. Combined with the powerful parallel computing capabilities of the GPU, the present invention can efficiently train deep learning models, thereby obtaining high-precision ophthalmic disease classification results in a relatively short time.
[0072] The content has applicability for clinical application. The S3FNet model of the present invention can be widely used in various ophthalmic medical institutions, including hospitals and clinics, to assist doctors in early diagnosis of ophthalmic diseases. The high accuracy and robustness of the model enable it to run stably on image data of different devices and different patients, and has strong clinical adaptability. The promotion of this technology can help more medical institutions realize intelligent ophthalmic disease screening, reduce manual intervention, and improve overall diagnosis and treatment efficiency.
[0073] 4. Implementation of the training process
[0074] The training process of the present invention adopts a standard deep learning training process. The specific implementation steps are as follows:
[0075] Dataset preparation: Collect fundus medical images from different hospitals to construct training sets, validation sets, and test sets. The dataset includes a variety of ophthalmic diseases (such as diabetic retinopathy, glaucoma, etc.) and normal fundus images.
[0076] Data preprocessing: Perform color normalization, adaptive histogram equalization (CLAHE) and other preprocessing on the original image to enhance the contrast of the image and improve the details. Use data enhancement techniques (such as rotation, scaling, flipping, etc.) to expand the training set and increase the robustness of the model.
[0077] Model training: The S3FNet model is trained using the training set, using the Adam optimizer, and cross-validation to optimize hyperparameters (such as learning rate, batch size, etc.). To avoid overfitting, L2 regularization, early stopping strategy, and learning rate scheduling are used.
[0078] Model evaluation: Use the validation set to evaluate the trained model and calculate indicators such as accuracy, F1 score, AUC, etc. to ensure the classification performance of the model. Finally, use the test set to perform a final evaluation on the model to verify its performance on unknown data.
[0079] In order to verify the effectiveness of the ophthalmic image classification method based on selective state space fusion proposed in this paper, a series of experiments were conducted. The experimental data covers a variety of ophthalmic diseases, including diabetes, glaucoma, cataracts, AMD (age-related macular degeneration), hypertension, myopia and other diseases / abnormalities, as well as normal fundus images. The experimental results are shown in Table 1, showing the performance indicators of the S3FNet model on different ophthalmic disease classification tasks.
[0080] Table 1
[0081]
[0082] As can be seen from the table, the S3FNet model has achieved high classification performance in most disease image categories, especially in the classification of myopia and normal fundus images, with F1-Score reaching 0.9549738 and 0.9849746 respectively. Even in the classification of AMD and glaucoma, which are more difficult to classify, the model also showed good performance, with F1-Score of 0.8042236 and 0.8647399 respectively. The overall average F1-Score is 0.8949721, indicating that the model has high classification accuracy and robustness.
[0083] In summary, the present invention effectively improves the analysis accuracy and robustness of fundus medical images by innovatively combining wavelet multi-scale feature extraction, deep learning technology and visual-language joint learning. This method can not only fully exploit the multi-scale frequency domain features in fundus images, but also capture global information through the Transformer encoder, and combine the CLIP model for visual-language fusion learning, which significantly improves the classification accuracy of ophthalmic disease images.
Claims
1. A method for ophthalmic image classification based on selective state space fusion, characterized in that: The steps include: Step 1: Obtain fundus medical image data, perform preprocessing, and generate a fundus medical image data set; Step 2: Build the S3FNet model and realize the classification of fundus medical images through the collaborative work of wavelet multi-scale feature extractor, inductive bias Transformer encoder IBTE and multi-model aggregation module MA; Step 3: Use the training set of fundus medical imaging dataset to train the S3FNet model and optimize it by adjusting the network parameters; Step 4: Use the test set in the fundus medical image dataset to evaluate the trained S3FNet model and generate the final classification model.
2. The ophthalmic image classification method based on selective state space fusion according to claim 1, characterized in that: The preprocessing process in step 1 specifically includes: fundus medical image data collection and annotation; standardization processing to make the color value of the image conform to a unified range; adaptive histogram equalization to enhance the details in the image; rotation, scaling, flipping, cropping to expand the data set; and finally dividing the data set into a training set, a validation set, and a test set.
3. The ophthalmic image classification method based on selective state space fusion according to claim 2, characterized in that: The wavelet multi-scale feature extractor extracts multi-scale features of different frequencies from fundus medical images, and is implemented as follows: Perform wavelet transform on the input fundus medical image, use multi-level decomposition to extract different frequency information in the image, and obtain feature maps of multiple scales through convolutional network CNN; The formed feature maps of multiple scales are input into the Mamba architecture based on the selective state space model SSM to optimize the representation of features and enhance the distinguishability and expressiveness of features; The features processed by the Mamba architecture are spliced with the fundus medical images input before processing to generate a multi-level feature map.
4. The ophthalmic image classification method based on selective state space fusion according to claim 3, characterized in that: The inductive biased Transformer encoder IBTE fuses local features and global features through a cross-attention mechanism, which is implemented as follows: The input fundus medical image is decomposed into multiple small blocks by block division, and each small block is converted into a feature vector; For each feature vector, the spatial position information is enhanced using the position encoding information to form a query sequence; The query sequence is cross-attended with the multi-level feature maps output by the wavelet multi-scale extractor to obtain global features. The feature maps of multiple scales extracted by CNN and the global features learned by Transformer are concatenated. The global context modeling of the input fundus medical image is performed through a multi-layer Transformer module to extract the global visual features of the image.
5. The ophthalmic image classification method based on selective state space fusion according to claim 4, characterized in that: The multi-model aggregation module MA combines vision-language joint learning to improve the classification performance of ophthalmic diseases, achieving the following: Use CLIP, a contrastive language-image pre-training model, to fuse the visual features of fundus medical images with the corresponding text descriptions; Through contrastive learning, the similarity between visual features and corresponding text descriptions is calculated, so that the model can understand the relationship between images and texts, and generate the final classification probability through the Softmax normalization method.
Citation Information
Cited By
Non-contact blood pressure detection method and system based on deep learning
CN120336833A
Medical image classification method and system based on multi-scale spatial state modeling
CN121147641A
Medical image classification method and system based on multi-scale spatial state modeling
CN121147641B