Interactive attention-based eye fundus image processing method, system and terminal
By constructing a fundus image processing model that includes bidirectional state space branches, CNN branches and interactive attention fusion modules, the problem of inability to effectively capture long-distance dependencies in the prior art is solved, which improves the classification accuracy of fundus images and reduces the computational overhead.
Patent Information
- Application Number
- CN202411909360.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, fundus image processing based on interactive attention is limited by the small receptive field of convolutional operations and cannot effectively capture long-distance dependencies, resulting in insufficient accuracy in fundus image classification and high computational overhead.
The fundus image processing model is constructed using bidirectional state space branches, CNN branches, interactive attention fusion module (IAFM) and multi-layer perceptrons. Global features are extracted through bidirectional state space branches, local features are extracted by CNN branches, and feature fusion and classification are performed through IAFM module.
Improves the classification accuracy of fundus images, reduces computational overhead, and enhances feature interaction and attention integration.
Smart Images

Figure CN119963879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image data processing, and in particular to a fundus image processing method, system, terminal and computer-readable storage medium based on interactive attention. Background Art
[0002] At present, many deep learning techniques have shown good results in fundus image processing based on interactive attention. For example, some people have introduced maximum and average aggregation operators and multi-scale information fusion methods, used transfer learning to train CNN (Convolutional Neural Networks), and further used temporal position encoding and time-sensitive self-attention to predict glaucoma from irregularly sampled fundus images. Others have designed a dual-branch network that integrates the features of ViT (Vision Transformer) and CNN branches.
[0003] However, existing fundus image analysis methods are often limited by the small receptive field of convolution operations, making it challenging to capture long-range dependencies. Although the Transformer model can alleviate some of these limitations, its complex computational requirements remain a significant constraint, resulting in current deep network models that are inaccurate and computationally expensive for fundus image classification.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a fundus image processing method, system, terminal and computer-readable storage medium based on interactive attention, aiming to solve the problem that the fundus image processing technology based on interactive attention in the prior art is limited by the small receptive field of convolution operation and cannot capture long-distance dependencies, resulting in inaccurate classification of fundus images and high computational overhead.
[0006] To achieve the above object, the present invention provides a fundus image processing method based on interactive attention, the fundus image processing method based on interactive attention comprising the following steps:
[0007] Acquire a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image;
[0008] Constructing a fundus image processing model, training and testing the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multilayer perceptron;
[0009] Inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features, and inputting the target fundus image into the CNN branch of the target model to extract features to obtain local features;
[0010] The global features and the local features are input into the IAFM module of the target model for feature fusion to obtain fused features, the fused features are downsampled, and the downsampled fused features are input into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
[0011] Optionally, the fundus image processing method based on interactive attention, wherein the training and testing of the fundus image processing model further comprises:
[0012] Acquiring historical data, the historical data including historical retinal images and classification results corresponding to the historical retinal images;
[0013] Using the historical retinal images as training samples and the classification results corresponding to the historical retinal images as labels to construct a data set;
[0014] The data set is divided into a training set, a test set and a validation set according to a preset ratio. The training set is used to train the fundus image processing model, the test set is used to evaluate the fundus image processing model in each round of training, and the validation set is used to evaluate the trained fundus image processing model.
[0015] Optionally, the fundus image processing method based on interactive attention, wherein the step of inputting the target fundus image into the bidirectional state space branch of the target model for feature extraction to obtain global features, further comprises:
[0016] Segmenting the target fundus image to obtain a plurality of image blocks of fixed size, and linearly embedding the plurality of image blocks into a high-dimensional space;
[0017] The plurality of image blocks in the high-dimensional space are segmented to obtain a global image for global feature extraction and a local image for local feature extraction.
[0018] Optionally, in the fundus image processing method based on interactive attention, the bidirectional state space branch includes a forward state space model, a backward state space model and a gated multilayer perceptron model.
[0019] Optionally, the fundus image processing method based on interactive attention, wherein the step of inputting the target fundus image into the bidirectional state space branch of the target model for feature extraction to obtain global features, specifically includes:
[0020] adding a position code to the global image to obtain data containing position information, and inputting the data containing the position information into the bidirectional state space branch for layer normalization to obtain normalized data;
[0021] Performing a convolution operation on the normalized data to obtain one-dimensional data, flipping the one-dimensional data to obtain a flipped image sequence, inputting the one-dimensional data into a backward state space model for reverse SSM scanning, and inputting the flipped image sequence into the forward state space model for forward SSM scanning, and outputting overall structural features and background features;
[0022] The overall structural features and the background features are spliced to obtain spliced features, the spliced features are input into the gated multi-layer perceptron model for integration, the integrated features are output, and the integrated features are matrix-added with the normalized data to obtain global features.
[0023] Optionally, in the fundus image processing method based on interactive attention, the CNN branch is a residual block structure in ResNet50, and the CNN branch is used to process the local image to extract local features.
[0024] Optionally, the fundus image processing method based on interactive attention, wherein the step of inputting the global feature and the local feature into the IAFM module of the target model for feature fusion to obtain the fused feature, specifically includes:
[0025] Inputting the global features and the local features into the IAFM module for convolution to obtain a first intermediate global feature and a first intermediate local feature;
[0026] Adding a SiLU activation function to the first intermediate global feature and the first intermediate local feature, and inputting the first intermediate global feature and the first intermediate local feature to which the SiLU activation function is added into a depthwise separable convolution to obtain a second intermediate global feature and a second intermediate local feature;
[0027] Adding a Sigmoid activation function to the second intermediate global feature and the second intermediate local feature to obtain a global weight and a local weight;
[0028] Performing a Hadamard product operation on the global weight and the local feature to obtain a first cross feature, and performing a Hadamard product operation on the local feature and the global weight to obtain a second cross feature;
[0029] The first cross feature and the second cross feature are connected, and the connected cross features are aggregated through the shuffle attention mechanism to obtain a fused feature.
[0030] In addition, to achieve the above-mentioned purpose, the present invention also provides a fundus image processing system based on interactive attention, wherein the fundus image processing system based on interactive attention comprises:
[0031] An image preprocessing module is used to obtain a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image;
[0032] A target model acquisition module is used to construct a fundus image processing model, train and test the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multi-layer perceptron;
[0033] A feature extraction module, used for inputting the target fundus image into the bidirectional state space branch of the target model to perform feature extraction to obtain global features, and inputting the target fundus image into the CNN branch of the target model to perform feature extraction to obtain local features;
[0034] A feature fusion classification module is used to input the global features and the local features into the IAFM module of the target model for feature fusion to obtain fused features, downsample the fused features, and input the downsampled fused features into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
[0035] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a fundus image processing program based on interactive attention stored on the memory and executable on the processor, wherein the fundus image processing program based on interactive attention, when executed by the processor, implements the steps of the fundus image processing method based on interactive attention as described above.
[0036] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a fundus image processing program based on interactive attention, and when the fundus image processing program based on interactive attention is executed by a processor, the steps of the fundus image processing method based on interactive attention as described above are implemented.
[0037] In the present invention, a fundus image to be processed is obtained, and the fundus image to be processed is preprocessed to obtain a target fundus image; a fundus image processing model is constructed, and the fundus image processing model is trained and tested to obtain a target model, and the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multi-layer perceptron; the target fundus image is input into the bidirectional state space branch of the target model for feature extraction to obtain a global feature, and the target fundus image is input into the CNN branch of the target model for feature extraction to obtain a local feature; the global feature and the local feature are input into the IAFM module of the target model for feature fusion to obtain a fusion feature, the fusion feature is downsampled, and the downsampled fusion feature is input into the multi-layer perceptron of the target model for classification to obtain a classification result of the fundus image to be processed. The present invention enhances feature interaction and promotes attention-based integration between two branches through interactive attention fusion, thereby improving the accuracy of fundus image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flow chart of a preferred embodiment of the fundus image processing method based on interactive attention of the present invention;
[0039] Figure 2 It is a structural diagram of a fundus image processing model in the fundus image processing method based on interactive attention of the present invention;
[0040] Figure 3 It is a structural diagram of the IAFM module in the fundus image processing method based on interactive attention of the present invention;
[0041] Figure 4 It is a structural diagram of a preferred embodiment of the fundus image processing system based on interactive attention of the present invention;
[0042] Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0043] The present application provides a fundus image processing method, system and terminal based on interactive attention. In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0045] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0046] The fundus image processing method based on interactive attention described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the fundus image processing method based on interactive attention includes the following steps:
[0047] Step S10: Acquire a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image.
[0048] The fundus usually refers to the inside of the back of the eyeball, which is composed of the optic disc, macula, retina and some blood vessels. It appears orange-red and can usually be explored through imaging technology. In this embodiment, the fundus image of the target user is acquired through a fundus camera or other equipment, and the fundus image is preprocessed to achieve the purpose of data enhancement. The preprocessing process includes random cropping (random cropping is an operation of randomly selecting an area of the input image for cropping during the training process, aiming to generate multiple different training samples, thereby increasing the diversity of the data and improving the generalization ability and robustness of the model), horizontal flipping, random rotation and other data enhancement technologies, and finally obtains the target fundus image for input model.
[0049] Step S20, constructing a fundus image processing model, training and testing the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multilayer perceptron.
[0050] It is understandable that the present invention proposes a fundus image processing model for processing fundus images, and the fundus image processing model uses the CNN-Mamba interactive fusion network framework (CMIFNet). The network consists of a local feature extraction branch based on CNN and a global information integration branch based on a bidirectional state space model, and an interactive attention fusion module is designed to enhance feature interaction and promote attention-based integration between the two branches. Through experiments, the method of the present invention shows good recognition performance on clinical data sets.
[0051] Specifically, construct Figure 2 The fundus image processing model shown in the figure includes a bidirectional state space model (BSSM), a CNN branch and an interactive attention fusion module (IAFM). The BSSM branch consists of two SSM pathways, and the BSSM branch extracts global information of the fundus image from two directions, including the overall structure of the retina and background information. The CNN branch focuses on local features, such as morphological changes of new blood vessels and retinal hemorrhages in specific locations. The IAFM module is designed to fuse global and local information. Finally, a multilayer perceptron is used as the classifier head to obtain the final classification result.
[0052] Furthermore, the training and testing of the fundus image processing model may also include:
[0053] Acquiring historical data, the historical data including historical retinal images and classification results corresponding to the historical retinal images;
[0054] Using the historical retinal images as training samples and the classification results corresponding to the historical retinal images as labels to construct a data set;
[0055] The data set is divided into a training set, a test set and a validation set according to a preset ratio. The training set is used to train the fundus image processing model, the test set is used to evaluate the fundus image processing model in each round of training, and the validation set is used to evaluate the trained fundus image processing model.
[0056] In this embodiment, a historical retinal image and the classification result corresponding to the historical retinal image are obtained, the historical retinal image is used as a training sample, and the classification result corresponding to the historical retinal image is used as the corresponding true label to construct a data set. The pre-processed data is divided into a training set, a test set, and a validation set according to a preset ratio (for example, the data ratio of the training set, the test set, and the validation set is 6:2:2), the training set is used to train the fundus image processing model and optimize the parameters of the model, and the test set is used to detect the accuracy of the prediction results of the trained fundus image processing model, thereby evaluating the performance of the model.
[0057] Furthermore, the fundus image processing model is trained and tested to obtain a target model. In order to save computing resources, the input image is adjusted to 224×224 in this embodiment. The learning rate is set to 0.0001 and the batch size is set to 32. This embodiment uses the Adam (adaptive moment estimation) optimizer, and the first-order and second-order momentum decay values are 0.9 and 0.999, respectively. Cross entropy is used as the loss function. The experiments were conducted on two NVIDIATITAN XP GPUs, each with 12GB of memory. This embodiment uses a 5-fold cross validation method for comparison, which helps to alleviate the overfitting of the model.
[0058] Training and testing the fundus image processing model is a binary classification task, in which the evaluation indicators include accuracy, precision, F1-score and area under the ROC curve (AUC) as evaluation indicators to evaluate and compare the performance of the model, and the calculation method is as follows:
[0059]
[0060]
[0061] Among them, TP (true positive), TN (true negative), FP (false positive) and FN (false negative) are the number of true positive, true negative, false positive and false negative samples respectively, TPR and FPR represent the true positive rate and false positive rate, and Recall represents the recall rate.
[0062] Step S30, inputting the target fundus image into the bidirectional state space branch of the target model for feature extraction to obtain global features, and inputting the target fundus image into the CNN branch of the target model for feature extraction to obtain local features.
[0063] It can be understood that the step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features and obtain global features may also include:
[0064] Segmenting the target fundus image to obtain a plurality of image blocks of fixed size, and linearly embedding the plurality of image blocks into a high-dimensional space;
[0065] The plurality of image blocks in the high-dimensional space are segmented to obtain a global image for global feature extraction and a local image for local feature extraction.
[0066] Furthermore, the bidirectional state space branch includes a forward state space model, a backward state space model and a gated multilayer perceptron model.
[0067] The step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features specifically includes:
[0068] adding a position code to the global image to obtain data containing position information, and inputting the data containing the position information into the bidirectional state space branch for layer normalization to obtain normalized data;
[0069] Performing a convolution operation on the normalized data to obtain one-dimensional data, flipping the one-dimensional data to obtain a flipped image sequence, inputting the one-dimensional data into a backward state space model for reverse SSM scanning, and inputting the flipped image sequence into the forward state space model for forward SSM scanning, and outputting overall structural features and background features;
[0070] The overall structural features and the background features are spliced to obtain spliced features, the spliced features are input into the gated multi-layer perceptron model for integration, the integrated features are output, and the integrated features are matrix-added with the normalized data to obtain global features.
[0071] It is understandable that the state space model (SSM) was originally designed to process one-dimensional dense data in natural language processing (NLP) and cannot capture the spatial features of the image. To address this limitation, the present embodiment uses the SSM bidirectional scanning method to compress the image block into one-dimensional data. The image sequence is flipped to perform forward and reverse SSM scans, so that the spatial context features can be perceived. In this way, the BSSM branch can naturally capture long-distance dependencies while retaining the integrity of the information. In order to better integrate the features extracted from two directions, the present embodiment introduces a gated multi-layer perceptron model after the BSSM. This design uses a multi-layer perceptron and a gating mechanism to extract spatial context from the image. In addition, before the image features are input into the BSSM, the present embodiment adds position encoding to the one-dimensional image tag, enabling the model to learn spatial sequence information. This improves its ability to capture directional clues and effectively utilize image context features, thereby more sensitively grasping the global state of the image and better identifying the lesion site.
[0072] Furthermore, the target fundus image is input into the CNN branch of the target model for feature extraction to obtain local features. The CNN branch is a residual block structure in ResNet50, and the CNN branch is used to process the local image to extract local features.
[0073] Specifically, the CNN branches from left to right are BN (BatchNorm), Conv1×1, BN, Conv3×3, BN, Conv1×1. For the CNN branch, the segmented slices are processed using the residual block in ResNet50 to extract local information between slice blocks. Compared with applying convolution to the entire image, this method reduces the amount of computation, reduces the dependence on channel complexity, and can extract more advanced semantic features.
[0074] Step S40, input the global features and the local features into the IAFM module of the target model for feature fusion to obtain fused features, downsample the fused features, and input the downsampled fused features into the multilayer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
[0075] The step of inputting the global features and the local features into the IAFM module of the target model to perform feature fusion to obtain fused features specifically includes:
[0076] Inputting the global features and the local features into the IAFM module for convolution to obtain a first intermediate global feature and a first intermediate local feature;
[0077] Adding a SiLU activation function to the first intermediate global feature and the first intermediate local feature, and inputting the first intermediate global feature and the first intermediate local feature to which the SiLU activation function is added into a depthwise separable convolution to obtain a second intermediate global feature and a second intermediate local feature;
[0078] Adding a Sigmoid activation function to the second intermediate global feature and the second intermediate local feature to obtain a global weight and a local weight;
[0079] For the global weight and the local feature F l Performing a Hadamard product operation to obtain a first cross feature, and performing a Hadamard product operation on the local feature and the global weight to obtain a second cross feature;
[0080] The first cross feature and the second cross feature are connected, and the connected cross features are aggregated through the shuffle attention mechanism to obtain a fused feature.
[0081] It is understandable that simple connection is not enough to effectively fuse the features captured by the BSSM branch and the CNN branch. In order to enhance the feature fusion between the two branches, this embodiment designs an interactive attention feature fusion module for feature interaction and integration. Specifically, Figure 3 As shown, this embodiment first extracts the local features F from the CNN and BSSM branches l and the global feature F g Input into the depth-wise separable convolution and then input into the activation function to obtain the corresponding weight W g and W l , the calculation process can be expressed as follows:
[0082] W g =δ(DSConv(SiLU(Conv(F g ))));
[0083] W l =δ(DSConv(SiLU(Conv(F l ))));
[0084] Among them, Conv(·) represents the convolution operation, SiLU(·) is the SiLU activation function, δ is the Sigmoid function, and DSConv(·) represents the depthwise separable convolution.
[0085] Furthermore, simple gating is used to reduce computational overhead. A simple gating mechanism is implemented through matrix dot product to cross-map global and local features to corresponding branches. Finally, the global and local feature maps are applied to their respective branches through channel connection and shuffle attention mechanism to complete the aggregation of feature maps. The process can be described as follows:
[0086]
[0087] Among them, F f represents the final output of IAFM (i.e., fusion feature), and Cat(·) represents the gated weighted feature F l and F g , SA(·) represents the shuffle attention mechanism, Represents the Hadamard product operation.
[0088] Furthermore, the fused features are downsampled, and after being downsampled, the fused features are subsequently passed to the next CNN and BSSM modules to reduce the complexity of the model. After passing through all CNN and BSSM modules, the final features are input into a multi-layer perceptron for classification to obtain the classification results of the fundus image to be processed.
[0089] Furthermore, if Figure 4 As shown, based on the above-mentioned fundus image processing method based on interactive attention, the present invention also provides a fundus image processing system based on interactive attention, wherein the fundus image processing system based on interactive attention includes:
[0090] An image preprocessing module 51 is used to obtain a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image;
[0091] A target model acquisition module 52 is used to construct a fundus image processing model, train and test the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multi-layer perceptron;
[0092] A feature extraction module 53, used for inputting the target fundus image into the bidirectional state space branch of the target model to perform feature extraction to obtain global features, and inputting the target fundus image into the CNN branch of the target model to perform feature extraction to obtain local features;
[0093] The feature fusion classification module 54 is used to input the global features and the local features into the IAFM module of the target model for feature fusion to obtain fused features, downsample the fused features, and input the downsampled fused features into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
[0094] Furthermore, if Figure 5 As shown, based on the above-mentioned interactive attention-based fundus image processing method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0095] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a fundus image processing program 40 based on interactive attention is stored on the memory 20, and the fundus image processing program 40 based on interactive attention can be executed by the processor 10, thereby realizing the fundus image processing method based on interactive attention in the present application.
[0096] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the interactive attention-based fundus image processing method.
[0097] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0098] In one embodiment, when the processor 10 executes the interactive attention-based fundus image processing program 40 in the memory 20, the following steps are implemented:
[0099] Acquire a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image;
[0100] Constructing a fundus image processing model, training and testing the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multilayer perceptron;
[0101] Inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features, and inputting the target fundus image into the CNN branch of the target model to extract features to obtain local features;
[0102] The global features and the local features are input into the IAFM module of the target model for feature fusion to obtain fused features, the fused features are downsampled, and the downsampled fused features are input into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
[0103] The training and testing of the fundus image processing model may also include:
[0104] Acquiring historical data, the historical data including historical retinal images and classification results corresponding to the historical retinal images;
[0105] Using the historical retinal images as training samples and the classification results corresponding to the historical retinal images as labels to construct a data set;
[0106] The data set is divided into a training set, a test set and a validation set according to a preset ratio. The training set is used to train the fundus image processing model, the test set is used to evaluate the fundus image processing model in each round of training, and the validation set is used to evaluate the trained fundus image processing model.
[0107] The step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features may also include:
[0108] Segmenting the target fundus image to obtain a plurality of image blocks of fixed size, and linearly embedding the plurality of image blocks into a high-dimensional space;
[0109] The plurality of image blocks in the high-dimensional space are segmented to obtain a global image for global feature extraction and a local image for local feature extraction.
[0110] The bidirectional state space branch includes a forward state space model, a backward state space model and a gated multilayer perceptron model.
[0111] The step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features specifically includes:
[0112] adding a position code to the global image to obtain data containing position information, and inputting the data containing the position information into the bidirectional state space branch for layer normalization to obtain normalized data;
[0113] Performing a convolution operation on the normalized data to obtain one-dimensional data, flipping the one-dimensional data to obtain a flipped image sequence, inputting the one-dimensional data into a backward state space model for reverse SSM scanning, and inputting the flipped image sequence into the forward state space model for forward SSM scanning, and outputting overall structural features and background features;
[0114] The overall structural features and the background features are spliced to obtain spliced features, the spliced features are input into the gated multi-layer perceptron model for integration, the integrated features are output, and the integrated features are matrix-added with the normalized data to obtain global features.
[0115] Among them, the CNN branch is a residual block structure in ResNet50, and the CNN branch is used to process the local image to extract local features.
[0116] The step of inputting the global features and the local features into the IAFM module of the target model to perform feature fusion to obtain fused features specifically includes:
[0117] Inputting the global features and the local features into the IAFM module for convolution to obtain a first intermediate global feature and a first intermediate local feature;
[0118] Adding a SiLU activation function to the first intermediate global feature and the first intermediate local feature, and inputting the first intermediate global feature and the first intermediate local feature to which the SiLU activation function is added into a depthwise separable convolution to obtain a second intermediate global feature and a second intermediate local feature;
[0119] Adding a Sigmoid activation function to the second intermediate global feature and the second intermediate local feature to obtain a global weight and a local weight;
[0120] Performing a Hadamard product operation on the global weight and the local feature to obtain a first cross feature, and performing a Hadamard product operation on the local feature and the global weight to obtain a second cross feature;
[0121] The first cross feature and the second cross feature are connected, and the connected cross features are aggregated through the shuffle attention mechanism to obtain a fused feature.
[0122] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a fundus image processing program based on interactive attention, and when the fundus image processing program based on interactive attention is executed by a processor, the steps of the fundus image processing method based on interactive attention as described above are implemented.
[0123] In summary, the present invention proposes a fundus image processing method, system and terminal based on interactive attention, the method comprising: obtaining a fundus image to be processed, preprocessing the fundus image to be processed, and obtaining a target fundus image; constructing a fundus image processing model, training and testing the fundus image processing model, and obtaining a target model, the target model comprising: a bidirectional state space branch, a CNN branch, an IAFM module and a multi-layer perceptron; inputting the target fundus image into the bidirectional state space branch of the target model for feature extraction to obtain a global feature, inputting the target fundus image into the CNN branch of the target model for feature extraction to obtain a local feature; inputting the global feature and the local feature into the IAFM module of the target model for feature fusion to obtain a fusion feature, downsampling the fusion feature, and inputting the downsampled fusion feature into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed. The present invention enhances feature interaction and promotes attention-based integration between two branches through interactive attention fusion, thereby improving the accuracy of fundus image classification and having low computational overhead.
[0124] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.
[0125] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0126] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A fundus image processing method based on interactive attention, characterized in that: The fundus image processing method based on interactive attention includes: Acquire a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image; Constructing a fundus image processing model, training and testing the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multilayer perceptron; Inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features, and inputting the target fundus image into the CNN branch of the target model to extract features to obtain local features; The global features and the local features are input into the IAFM module of the target model for feature fusion to obtain fused features, the fused features are downsampled, and the downsampled fused features are input into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
2. The fundus image processing method based on interactive attention according to claim 1, characterized in that: The training and testing of the fundus image processing model also includes: Acquiring historical data, the historical data including historical retinal images and classification results corresponding to the historical retinal images; Using the historical retinal images as training samples and the classification results corresponding to the historical retinal images as labels to construct a data set; The data set is divided into a training set, a test set and a validation set according to a preset ratio. The training set is used to train the fundus image processing model, the test set is used to evaluate the fundus image processing model in each round of training, and the validation set is used to evaluate the trained fundus image processing model.
3. The fundus image processing method based on interactive attention according to claim 2, characterized in that: The step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features may also include: Segmenting the target fundus image to obtain a plurality of image blocks of fixed size, and linearly embedding the plurality of image blocks into a high-dimensional space; The plurality of image blocks in the high-dimensional space are segmented to obtain a global image for global feature extraction and a local image for local feature extraction.
4. The fundus image processing method based on interactive attention according to claim 3 is characterized in that: The bidirectional state space branch includes a forward state space model, a backward state space model and a gated multilayer perceptron model.
5. The fundus image processing method based on interactive attention according to claim 4 is characterized in that: The step of inputting the target fundus image into the bidirectional state space branch of the target model to extract features to obtain global features specifically includes: adding a position code to the global image to obtain data containing position information, and inputting the data containing the position information into the bidirectional state space branch for layer normalization to obtain normalized data; Performing a convolution operation on the normalized data to obtain one-dimensional data, flipping the one-dimensional data to obtain a flipped image sequence, inputting the one-dimensional data into a backward state space model for reverse SSM scanning, and inputting the flipped image sequence into the forward state space model for forward SSM scanning, and outputting overall structural features and background features; The overall structural features and the background features are spliced to obtain spliced features, the spliced features are input into the gated multi-layer perceptron model for integration, the integrated features are output, and the integrated features are matrix-added with the normalized data to obtain global features.
6. The fundus image processing method based on interactive attention according to claim 3 is characterized in that: The CNN branch is a residual block structure in ResNet50, and the CNN branch is used to process the local image to extract local features.
7. The fundus image processing method based on interactive attention according to claim 4 is characterized in that: The step of inputting the global features and the local features into the IAFM module of the target model to perform feature fusion to obtain fused features specifically includes: Inputting the global features and the local features into the IAFM module for convolution to obtain a first intermediate global feature and a first intermediate local feature; Adding a SiLU activation function to the first intermediate global feature and the first intermediate local feature, and inputting the first intermediate global feature and the first intermediate local feature to which the SiLU activation function is added into a depthwise separable convolution to obtain a second intermediate global feature and a second intermediate local feature; Adding a Sigmoid activation function to the second intermediate global feature and the second intermediate local feature to obtain a global weight and a local weight; Performing a Hadamard product operation on the global weight and the local feature to obtain a first cross feature, and performing a Hadamard product operation on the local feature and the global weight to obtain a second cross feature; The first cross feature and the second cross feature are connected, and the connected cross features are aggregated through the shuffle attention mechanism to obtain a fused feature.
8. A fundus image processing system based on interactive attention, characterized in that: The fundus image processing system based on interactive attention includes: An image preprocessing module is used to obtain a fundus image to be processed, and preprocess the fundus image to be processed to obtain a target fundus image; A target model acquisition module is used to construct a fundus image processing model, train and test the fundus image processing model to obtain a target model, wherein the target model includes: a bidirectional state space branch, a CNN branch, an IAFM module and a multi-layer perceptron; A feature extraction module, used for inputting the target fundus image into the bidirectional state space branch of the target model to perform feature extraction to obtain global features, and inputting the target fundus image into the CNN branch of the target model to perform feature extraction to obtain local features; A feature fusion classification module is used to input the global features and the local features into the IAFM module of the target model for feature fusion to obtain fused features, downsample the fused features, and input the downsampled fused features into the multi-layer perceptron of the target model for classification to obtain the classification result of the fundus image to be processed.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an interactive attention-based fundus image processing program stored in the memory and executable on the processor. When the interactive attention-based fundus image processing program is executed by the processor, the steps of the interactive attention-based fundus image processing method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a fundus image processing program based on interactive attention, and when the fundus image processing program based on interactive attention is executed by a processor, the steps of the fundus image processing method based on interactive attention as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Heterogeneous double-flow fusion method and system for grading diabetic retinopathy
CN121033041A
Heterogeneous dual-stream fusion method and system for diabetic retinopathy grading
CN121033041B
Pedestrian crossing intention prediction method based on man-vehicle interaction modeling and multi-modal fusion
CN121330653A
Texture image classification method based on Vision Mama double-branch cross fusion network
CN122265745A
A texture image classification method based on VisionMamba dual-branch cross-fusion network
CN122265745B