A modeling method of MEA-Net network based on multi-expert system
The MEA-Net network's multi-expert system solves the problems of difficult feature similarity differentiation and macular structure interference in traditional retinal OCT image classification, achieving efficient feature extraction and classification of macular edema and degeneration, and improving the model's accuracy and robustness.
Patent Information
- Application Number
- CN202211467595.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-11-22
AI Technical Summary
Traditional retinal OCT image classification methods suffer from difficulties in distinguishing feature similarities and interference from other macular structures when dealing with fine-grained classification problems in the medical field, resulting in low classification accuracy. Furthermore, existing two-dimensional image classification algorithms are not accurate enough in the analysis of macular edema and degeneration.
The MEA-Net network, based on a multi-expert system, is adopted. By combining input modules, voting modules, expert network modules, and output modules, the multi-expert network learns different features and strengthens feature analysis through a separate channel adjacency module. Combined with global average pooling and batch normalization layers, the accuracy and robustness of the model are improved.
It improves the accuracy of analysis and classification of macular edema and degeneration manifestations, enhances the model's generalization performance and anti-overfitting ability, and achieves efficient feature extraction and classification of OCT images.
Smart Images

Figure CN115861195B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision and medical imaging, and in particular to a modeling method of MEA-Net network based on a multi-expert system. BACKGROUND
[0002] The macula is located in the optic disc of the fundus of the eye, in the optical center area of the human eye, and is the projection point of the visual axis. The macula is the most concentrated part of the central visual cells of the human retina. The central depression of the macula is called the fovea, which is the most acute place of vision. Macular edema refers to the inflammatory reaction and fluid infiltration in the macular area of the fundus of the retina, resulting in fluid accumulation in the inner layer of the macular retina, causing severe visual loss. It is an ocular manifestation of diabetic retinopathy, central retinal vein occlusion and other eye diseases. Traditional fundus photography and eye B-ultrasound examination are time-consuming, complex to prepare, and the macular part is not clear, and cannot accurately reflect the macular lesion. Optical coherence tomography (Optical Coherence Tomography, abbreviated as OCT) uses the principle of light coherence to obtain the thickness and distance information provided by the reflection of different tissue interfaces in the eye and restores the image and data to perform tomographic scanning on the intraocular tissue. It can provide two-dimensional cross-sectional images or simulate three-dimensional stereoscopic images similar to pathological sections under a microscope to find very small macular lesions, greatly improving the visibility of the fundus; it is an important tool for tracking and following up on common diseases such as diabetic retinopathy, age-related macular degeneration, and retinal vein occlusion. It is an important means of detailed examination of the retina, especially the macular part and the optic disc.
[0003] Most of the traditional image classification methods use hand-crafted features or feature learning methods to describe the entire image, and then use a classifier to distinguish the object category. The accuracy of classification depends on the feature extraction method and quality, and improper selection can lead to low accuracy. Therefore, manual inspection and correction are often required, which is very tedious. With the development of deep learning, convolutional neural networks (Convolutional Neural Networks, CNN) have achieved excellent results in image classification. Currently, most existing retinal OCT image classification algorithms are based on two-dimensional image classification, but there are still some shortcomings in the fine-grained classification problem in the medical field. The first is that the image features of different positions of edema are similar and difficult to distinguish due to fewer features. The second is that other structures of the fundus macula can interfere with model analysis. SUMMARY
[0004] The present application provides a modeling method of MEA-Net network based on a multi-expert system.
[0005] A modeling method of MEA-Net network based on a multi-expert system, comprising the following steps:
[0006] S1: image preprocessing is performed on the original OCT image to obtain an OCT image set to be classified;
[0007] S2: a MEA-Net (Multi-Expert attention Net) network is constructed, and specifically includes:
[0008] an input module configured to perform preliminary analysis on the image;
[0009] a voting module configured to perform feature analysis on the preliminarily analyzed image and output voting input state weight P and image channel weight G;
[0010] an expert network module configured to receive the voting input state weight output by the voting module and perform deep analysis and processing on the features of the preliminarily analyzed image to obtain a feature map;
[0011] an output module configured to receive the image channel weight output by the voting module and combine the feature map output by the expert network module to obtain a vector outer product, and perform splicing in the channel dimension to output a tensor M, and then perform feature aggregation extraction and linear analysis to obtain a classification result;
[0012] The MEA-Net controls the input and output of the expert network, realizes feature extraction and analysis from different angles, and improves the accuracy and robustness of the model.
[0013] S3: the MEA-Net (Multi-Expert attention Net) network is trained using the OCT image set to be classified obtained in step 1), and a trained MEA-Net (Multi-Expert attention Net) network is obtained.
[0014] The present application includes data preprocessing, cropping, and proportionally distributing image data and label data sets; in order to improve the accuracy of fundus retinal lesion classification, the present application designs a Multi-Expert attention Net network model. The input module performs preliminary extraction on the data, the replaceable expert network module is responsible for deep analysis of the features, the voting module strengthens the attention of the features in the channel space through the separation channel adjacency module, the output weight realizes expert voting, and finally the output module obtains the classification result. The present application can use the picture information of optical coherence tomography, can analyze and extract the features of the four types of macular edema degeneration of the OCT image, PED, SRF, IRF, and HRF, and lays a foundation for improving the efficiency of subsequent retinal disease classification and analysis.
[0015] In step S1, the preprocessing includes:
[0016] The original OCT image is cut and adjusted to the pixel size, divided into a training set, a test set and a validation set, and image normalization processing is performed;
[0017] The different morphological images are standardized through the processing of the original OCT image, which facilitates the training and testing of the subsequent network model.
[0018] In step S2, the input module comprises a convolution layer, an average pooling layer and a relu function activation layer arranged in sequence.
[0019] The convolution and pooling layers are used to realize preliminary analysis of the image, so as to improve the effect of subsequent image feature analysis.
[0020] In step S2, the voting module comprises a weight generation unit and a separated channel adjacent module.
[0021] The weight generation unit comprises an MBConv module, a softmax layer and a full connection layer connected in sequence, the weight generation unit extracts features from the preliminary analyzed image through the MBConv module, realizes weight distribution through the softmax function, and obtains input state weight P. Wherein, the MBConv is a LinearBottleNecks layer with deep separable convolution and inverted residual.
[0022] The separated channel adjacent module comprises a convolution layer, a maximum pooling layer and four analysis units arranged in sequence.
[0023] Each of the four analysis units comprises a first convolution layer, a batch normalization layer, a maximum pooling layer, a second convolution layer and a spatial operation of a plane softmax layer arranged in sequence.
[0024] The separated channel adjacent module extracts and evaluates features from the preliminary analyzed image, and splices in the channel dimension to output a graph channel weight G.
[0025] Separating and strengthening the features in the channel dimension can extract features in multiple channels and improve the feature attention of the graph channel weight.
[0026] In step S2, the expert network module is a replaceable network composed of three peer expert networks, the expert network can be selected from mainstream neural networks such as Efficient-Net, Res-net, Dark-net and Dense-Net, or can be self-written, as long as the corresponding feature map output is output.
[0027] The deep analysis processing adopts the following formula:
[0028] Input=P iX (i = 1, 2, 3)
[0029] Wherein, Input is a feature map, P i is the i th weight tensor, E is the image of preliminary analysis.
[0030] The vector outer product is obtained, and specifically includes:
[0031]
[0032] Wherein, Out is the feature map obtained by the expert network module, G is the graph channel weight output by the voting module, j is the j th output module, k is the k th expert network output feature map, n is the channel dimension of the expert network output feature map, and M is an output tensor.
[0033] The deep analysis processing and vector channel dimension outer product are important methods for controlling the output feature map of the expert network module, and can improve the accuracy of the network model.
[0034] The output module includes an MBconv module, a batch normalization layer, a swish activation layer, a global average pooling layer, a drop random inactivation layer and a linear layer arranged in sequence.
[0035] The analyzed feature map is weighted and the feature is strengthened, the generalization performance of the network model is enhanced through the drop random inactivation, and finally the classification result is obtained through linear connection.
[0036] Compared with the prior art, the present application has the following advantages:
[0037] 1. The present application first proposes a MEA-Net network modeling method based on multi-expert system-MEA-Net (Multi-Expert attention Net), which can analyze and extract the features of retinal macular edema and degeneration of OCT images, and lay a foundation for improving the efficiency of subsequent retinal disease classification and analysis.
[0038] 2. The present application applies the idea of ensemble learning to the convolutional neural network for retinal macular lesion image classification, learns different features by using multi-expert network, weights the output of the classification task through the voting module, and adds the separated channel adjacency module to the voting module to strengthen the feature analysis ability of the voting module, so that it is easier to learn the commonality and difference between different classification tasks.
[0039] 3, The application proposes an attention mechanism based on channel space segmentation, and is named as a separated channel adjacent module. On one hand, the features are divided into 4 groups, and the features of each group of channel space are re-extracted and calibrated, so that the cross-channel integration information is obtained to enhance the original features; on the other hand, a batch normalization layer is added before attention enhancement, so that the training convergence speed of the separated channel adjacent module is strengthened, and the gradient disappearance problem is avoided.
[0040] 4, The output module of the application uses global average pooling, a batch normalization layer, and each output module corresponds to a classification label to generate a corresponding feature map, and the output modules are isolated from each other and have no interference; and the added drop layer can prevent the occurrence of network overfitting, and increase the robustness of the network. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 is the MEA-Net neural network model structure diagram in the application;
[0042] Figure 2 is the separated channel adjacent module structure diagram in the application;
[0043] Figure 3 is the voting module structure diagram in the application;
[0044] Figure 4 is the output module structure diagram in the application. DETAILED DESCRIPTION
[0045] The application will be further described below in combination with the drawings and specific implementation.
[0046] A MEA-Net network modeling method based on a multi-expert system, specifically comprising the following steps:
[0047] (1) Collect images and process and clean the data, and divide them into training, testing and validation sets. And the images are preprocessed.
[0048] In order to remove the original OCT images collected in medical treatment which contain fundus images and irrelevant information, unify the image size and reduce the calculation amount, the original image needs to be cropped. The specific method is: taking the upper right corner of the original image as the coordinate origin, and taking the image with a size of 384*384 pixels as the OCT image data, obtaining 3653 cleaned OCT data pictures, and each picture includes one or more classification information
[0049] In order to balance the data and prevent overfitting, data enhancement is performed on the several data sets, and the methods used include random horizontal flip, stretching transformation, Gaussian blur and translation transformation operation.
[0050] (2) Establishing the MEA-Net neural network model, selecting a suitable expert network model, and loading the corresponding pre-training weight.
[0051] As shown in Figure 1 MEA-Net is composed of an input module, a multi-expert network module, a voting module and an output module.
[0052] The input module is composed of a convolutional layer a1 (convolution kernel size 3X3, step 1, padding 1), an average pooling layer b1 (pooling kernel 3X3, step 1, padding 1) and a relu activation function layer (relu), which is a preliminary extraction of the input image features. The number of feature map channels remains unchanged in this process.
[0053] The multi-expert network is responsible for further analysis and extraction of feature maps, and is the core analysis part of the network. The multi-expert network is composed of three expert networks with the same structure. The selection of expert networks can be selected from mainstream networks, so that the network model can be adapted. The output feature map needs to be fixed to (120, 24, 24) size when used. Here, efficientnet_b2 is taken as an example, the network input tensor is (3, 384, 384), and the three network output feature maps are spliced and combined in the channel to form a feature fusion map tensor (360, 24, 24) as the output.
[0054] The voting module is composed of two parts, as shown in Figure 2 The features are selected and filtered as the selection input and evaluation weight of the expert network, which enhances the analysis, robustness and accuracy of the features. The first part is a separate channel adjacent module, as shown in Figure 3 The convolutional pooling layer is used to further refine the features. The analysis unit is composed of a convolutional layer a2.1 (convolution kernel size 1X1, step 1, padding 0), a relu activation function layer (relu), a batch normalization layer (BatchNormal), a maximum pooling layer (MaxPooling) b2.1 (pooling kernel size 3X3, step 1, padding 0), a convolutional layer b2.2 (convolution kernel size 1X1, step 1, padding 0). The second part is composed of an MBConv module (output channel size is 3) and a full connection layer (FullConnect). After the batch normalization layer fixes the mean and variance of the data, the feature map is mapped to a weight vector through the full connection layer, realizing the control of the expert network.
[0055] The voting module works as follows: (a) The feature map of the input layer is divided into four sub-tensors on each channel, and these sub-tensors are fed into the adjacency module of the separate channels to obtain the corresponding weight vectors and feature maps. (b) The four weight vectors are multiplied by the feature map of the input module to obtain the feature map input to the expert network, using the following formula:
[0056] Input i =P i ×E(i=1,2,3)
[0057] (c) Take the outer product of the three elements of the output feature maps from the three expert networks and the output tensor from the voting module, and concatenate them along the channel dimension, as shown in the formula:
[0058]
[0059] The output module analyzes, summarizes, and classifies the enhanced feature maps. For example... Figure 4 As shown, the output module consists of four sub-modules: the MBConv module, a batch normalization layer (BatchNormal), a global average pooling layer (MaxPooling) (with a 1x1 kernel size), a random dropout layer (probability 0.3), and a fully connected layer (FullConnect). Here, the MBConv module extracts features from the enhanced feature map, converts it into a linear tensor after batch normalization and global average pooling, and then maps it to a class tensor through the fully connected layer. The four labels are repeated four times. Finally, the four output tensors are concatenated and transformed to form the final output tensor P = {p1, p2, p3, p4}, representing the probabilities of each classification being true.
[0060] (3) Model training and testing. After data augmentation of the dataset, the model is trained. After training, the trained model is loaded, and the model is run on the test set to obtain various evaluation indicators and compared with mainstream models. The input data is used to obtain the predicted class vector based on the neural network model of this invention. The loss error between the predicted class vector and the true class vector is back-propagated to the neural network model to update and optimize the weights.
[0061] Specific training method: The entire dataset is divided into training, validation, and test sets in a 6:2:2 ratio. A suitable expert network is selected, and the corresponding pre-trained model is loaded. During training, the Adam optimization algorithm is used with the F1-Loss loss function. The overall network learning rate is fixed at 1e-4, and the minimum batch size is set to 4. The training is performed on an NVIDIA RTX 2080Ti graphics card for 400 epochs. Every 4 iterations, the learning strategy is adjusted based on the performance metrics of the validation set to obtain the final trained model, which is then tested on the test set.
[0062] In order to verify the effectiveness of the present application, 3653 lesion images are used to make a data set by implementing step (1) to train, extract the corresponding trained model to test, compare with multiple comparative experiments, and evaluate the classification evaluation results with accuracy (Accuracy), precision (precision), recall (recall) and F1-score. The network of the present application is compared with the original Efficient-net and Res-Net, and the comparison results are shown in Table 1.
[0063] Table 1 Experimental results
[0064] Network Accuracy Precision Recall F1-score MEA-Net (Efficientnet_b2 based) 0.995 0.994 0.986 0.995 Efficientnet_b2 0.993 0.995 0.990 0.993 MEA-Net (Res-Net18 based) 0.994 0.993 0.982 0.992 Res-Net 18 0.983 0.982 0.983 0.983
[0065] As shown in Table 1, the network of the MEA-Net model structure has different degrees of improvement in various parameter indicators compared with the corresponding single expert network model, and the performance is better than the original network. Compared with the original network, the present network increases the number of parameters and the expert network mechanism, so that the accuracy and robustness of the network are improved.
[0066] The embodiments of the present application described above do not constitute a limitation on the scope of protection of the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the claims of the present application.
Claims
1. A modeling method of MEA-Net network based on multi-expert system, characterized in that, The method comprises the following steps: S1: image preprocessing is performed on the original OCT image to obtain an OCT image set to be classified; S2: an MEA-Net network is constructed, specifically comprising: an input module for performing preliminary analysis on the image; a voting module for performing feature analysis on the preliminarily analyzed image and outputting a voting input state weight P and a graph channel weight G; an expert network module for receiving the voting input state weight output by the voting module and combining the preliminarily analyzed image to perform deep analysis processing on the features to obtain a feature map; An output module is configured to receive the channel weight of the voting output by the voting module, and to combine the vector outer product of the feature map output by the expert network module, and to splice in the channel dimension, and to output a tensor After feature aggregation extraction and linear analysis, a classification result is obtained. the deep analysis processing is processed by the following formula: ; wherein, is a feature map, is a first weight tensor, is a preliminary analyzed image; vector outer product, specifically comprising: ; wherein, is the feature map obtained by the expert network module, G is the channel weight of the output map of the voting module, j is the jth output module, k is the kth output feature map of the expert network, n is the channel dimension of the output feature map of the expert network, and M is the output tensor. S3: using the OCT image set to be classified obtained in step 1) to train the MEA-Net network to obtain a trained MEA-Net network.
2. The method of modeling of MEA-Net network based on multi-expert system according to claim 1, characterized in that, In step S1, the preprocessing comprises: cutting the original OCT image and adjusting it to a pixel size, dividing it into a training set, a test set and a verification set, and performing image normalization processing.
3. The method of modeling of MEA-Net network based on multi-expert system according to claim 1, characterized in that, In step S2, the input module comprises a convolution layer, an average pooling layer and a relu function activation layer arranged in sequence.
4. The method of modeling of MEA-Net network based on multi-expert system according to claim 1, characterized in that, In step S2, the voting module comprises a weight generation unit and a separate channel adjacent module.
5. The method of modeling of MEA-Net network based on multi-expert system according to claim 4, characterized in that, In step S2, the weight generation unit comprises an MBConv module, a softmax layer and a full connection layer connected in sequence, the weight generation unit extracts features from the preliminarily analyzed image through the MBConv module, and realizes weight distribution through the softmax function to obtain the input state weight P.
6. The method of modeling of MEA-Net network based on multi-expert system according to claim 4, characterized in that, In step S2, the separate channel adjacent module comprises a convolution layer, a maximum pooling layer and four analysis units arranged in sequence. Each analysis unit in the four analysis units comprises a first convolution layer, a batch normalization layer, a maximum pooling layer, a second convolution layer and a spatial operation of a plane softmax layer arranged in sequence.
7. The method of modeling of MEA-Net network based on multi-expert system according to claim 6, characterized in that, In step S2, the separate channel adjacent module extracts features from the preliminarily analyzed image and evaluates the features, and splices in the channel dimension to output the graph channel weight G.
8. The method of modeling of MEA-Net network based on multi-expert system according to claim 1, characterized in that, In step S2, the output module comprises an MBconv module, a batch normalization layer, a swish activation layer, a global average pooling layer, a drop random inactivation layer and a linear layer arranged in sequence.
Citation Information
Patent Citations
Medical image data analysis method and system based on fused deep tensor nerve network
CN107622485A
Pathological image classification device and method based on depth feature fusion and use method of device
CN114139588A