Brain tumor image classification transfer learning method based on multi-model combination
By constructing a parallel dual-branch structure of ResNet50 and EfficientNetB0, combined with data augmentation and transfer learning, the problems of low accuracy and high resources in brain tumor MRI image classification are solved, efficient and accurate classification effects are achieved, and the generalization ability and computing efficiency of the model are improved.
Patent Information
- Application Number
- CN202510594362.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
AI Technical Summary
There are problems in the existing brain tumor MRI image classification with low classification accuracy, high computing resource requirements, easy overfitting and insufficient generalization capabilities, especially under limited data sets and resource conditions.
Using a multi-model combination method, a parallel dual-branch structure of ResNet50 and EfficientNetB0 is constructed, the data set is expanded through data augmentation technology, and the multi-scale feature fusion strategy and L2 regularization and Dropout layer optimization model are used to train and fine-tune with transfer learning, and the Adamax optimizer and 50-fold cross-validation are used for evaluation.
It significantly improves the accuracy of brain tumor MRI image classification, enhances the generalization ability of the model, and maintains efficient computing performance under limited resource conditions, is suitable for different clinical scenarios, and reduces hardware requirements.
Smart Images

Figure CN120526201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image analysis and deep learning technology, and in particular to a migration learning method for brain tumor image classification based on multi-model combination. Background Art
[0002] Brain tumors, a disease that originates from abnormal cell proliferation in the human central nervous system, can be divided into two main categories: benign and malignant, based on their biological characteristics and clinical behavior. Benign brain tumors typically grow slowly and are less likely to spread, while malignant brain tumors grow rapidly and easily invade surrounding brain tissue, posing a serious threat to the patient's life and health. During clinical diagnosis and treatment, accurately identifying the type of brain tumor is crucial for developing timely, scientifically effective treatment plans and improving patient prognosis.
[0003] In traditional brain tumor diagnosis, magnetic resonance imaging (MRI), with its high-resolution soft tissue imaging capabilities, has become an important auxiliary tool for doctors. Radiologists observe and analyze MRI images to determine the presence, location, size, and nature of brain tumors. However, this manual diagnostic method has many limitations. First, the diagnostic results are highly dependent on the doctor's professional experience and subjective judgment, and consistency between different doctors is difficult to ensure. Second, manual diagnosis is inefficient, especially when processing large amounts of image data, which can easily cause visual fatigue, leading to misdiagnosis or missed diagnoses. More importantly, manual diagnosis faces enormous challenges in identifying early, subtle lesions, and early diagnosis is crucial to improving the survival rate of brain tumor patients.
[0004] With the rapid development of artificial intelligence (AI), traditional machine learning (ML) techniques and computer-aided diagnosis (CAD) systems are increasingly being applied to brain tumor classification. These methods typically require manual design and extraction of image features, such as shape and texture features. These features are then fed into classifiers such as support vector machines (SVMs) and random forests (RFs) for classification. While these methods have improved diagnostic efficiency to a certain extent, manual feature extraction requires specialized knowledge and significant human effort. Furthermore, the extracted features often fail to fully and accurately describe the essential information of the image, limiting the accuracy and reliability of the classification.
[0005] In recent years, deep learning techniques, particularly convolutional neural networks (CNNs), have made significant progress in medical image analysis. CNNs possess powerful automatic feature extraction capabilities, enabling them to directly learn valuable features from image data without the need for manual feature design and extraction, significantly improving the efficiency and accuracy of feature extraction. Numerous studies have attempted to apply CNNs to brain tumor classification, and by constructing CNN models of varying structures and scales, they have achieved significant improvements in brain tumor classification performance.
[0006] However, the practical application of CNNs in brain tumor classification still faces numerous challenges. First, in the field of medical imaging, obtaining large amounts of high-quality labeled data is often extremely difficult. The limited size of datasets can easily lead to overfitting of CNN models, resulting in models that perform well on the training set but have poor generalization capabilities in practical applications. Second, compared to traditional machine learning methods, CNNs are computationally more complex, and the training process requires powerful computing resources, particularly high-performance graphics processing units (GPUs). This, to a certain extent, limits their widespread application in resource-limited environments.
[0007] To overcome the aforementioned issues faced by CNNs in brain tumor classification, transfer learning (TL) technology has gradually become a research hotspot. By pre-training the model on large-scale general datasets and then fine-tuning the model for specific medical imaging tasks, transfer learning can achieve good performance with limited training samples and a shorter training time. Numerous researchers have conducted extensive research on brain tumor classification based on transfer learning, achieving considerable results by selecting different pre-training models and fine-tuning strategies.
[0008] Although transfer learning has shown certain advantages in brain tumor classification, there are still some urgent problems to be solved in this field. On the one hand, the traditional CNN model structure is relatively simple, and it is difficult to fully explore the complex semantic information and features in the image, resulting in a bottleneck in improving the classification accuracy. On the other hand, the training process of transfer learning still requires a large amount of computing power and high requirements for hardware resources, which increases the cost and difficulty of practical application. In addition, most existing studies are based on single or a few data sets for training and testing, and the generalization ability of the model lacks sufficient verification. Under different clinical scenarios and data distributions, the performance of the model may drop significantly. Therefore, developing an efficient, accurate and highly generalizable brain tumor classification method has become an important research direction in the current field of medical image analysis. Summary of the Invention
[0009] The present invention aims to provide a transfer learning method for brain tumor image classification based on multi-model combination, which can solve the problems of low classification accuracy, high computing resource requirements, easy overfitting and insufficient generalization ability in the existing brain tumor MRI image classification, and realize efficient and accurate brain tumor MRI image classification.
[0010] To achieve the above objectives, the present invention adopts the following technical means:
[0011] A multi-model combined transfer learning method for brain tumor image classification includes the following steps:
[0012] Image preprocessing: perform edge contour cropping and scaling on brain tumor MRI images to retain the main contour area of the brain;
[0013] Data enhancement: Using data enhancement technology, we rotate and flip the pre-processed images, expand and balance the original brain tumor MRI images, and obtain an enhanced dataset.
[0014] Building a combined model: A multi-scale feature fusion strategy is used to construct a hybrid architecture, ResNet50+EfficientNetB0. ResNet50 and EfficientNetB0 work together in a parallel dual-branch structure. The input image after image preprocessing is simultaneously fed into both branches for feature extraction. After being processed by the global average pooling layer, the extracted features are concatenated through the Concatenate layer. A Dense layer is added after the merging layer, and L2 regularization is introduced. A Dropout layer is added, and finally a SoftMax function is used for multi-classification processing.
[0015] Model training and fine-tuning: Using the enhanced dataset obtained in the data augmentation step, based on the idea of transfer learning and the pre-trained model parameters, the constructed combined model is trained and fine-tuned for the brain tumor MRI image classification task;
[0016] Finally, the model is evaluated: 5-fold cross-validation is used on the Kaggle dataset to test the combined model after model training and fine-tuning, and the classification performance is compared and analyzed with other models and existing popular methods.
[0017] A further solution of the present invention is that the ResNet50 model is composed of multiple residual modules connected in series, each residual module includes a "bottleneck structure" consisting of a 1×1 convolutional layer, a 3×3 convolutional layer and a 1×1 convolutional layer, as well as a batch normalization layer and a ReLU activation function.
[0018] A further solution of the present invention is that the EfficientNetB0 model consists of 9 stages, of which the second to eighth stages are formed by stacking 16 MBConv modules, and the MBConv modules sequentially include 1×1 convolution for dimensionality increase, K×K depthwise separable convolution, SE attention module, 1×1 convolution for dimensionality reduction, Dropout layer and residual connection.
[0019] A further solution of the present invention is that in the parallel dual-branch structure, the features extracted by the ResNet50 branch and the EfficientNetB0 branch are processed by the global average pooling layer and then spliced and fused according to the channel dimension.
[0020] A further solution of the present invention is that during the model training and fine-tuning process, an Adamax optimizer is used, the learning rate is set to 0.0001-0.001, and the training rounds are set to 30-80 rounds.
[0021] A further solution of the present invention is that in the model evaluation step, in addition to accuracy, sensitivity, specificity, and AUC value indicators are also used to evaluate the performance of the combined model.
[0022] A further solution of the present invention is that when constructing a combined model, the ResNet50+EfficientNetV2B0 or EfficientNetB0+EfficientNetV2B0 architecture can also be selected to achieve brain tumor MRI image classification through the feature combination of the corresponding model branches.
[0023] Beneficial effects of the present invention:
[0024] 1. Significantly improve classification accuracy: By combining the powerful deep feature extraction capabilities of ResNet50 with the efficient computing power of EfficientNetB0, a hybrid architecture is constructed using a multi-scale feature fusion strategy. The residual connection mechanism of ResNet50 can effectively solve the gradient vanishing and degradation problems of deep networks, thereby learning rich high-level semantic information of the image; EfficientNetB0's compound scaling method and advanced MBConv module design enable it to reduce the amount of computation while maintaining performance. The two work together in a parallel dual-branch structure and fuse features. Compared with a single model, they can learn brain tumor MRI image features more comprehensively and deeply, significantly improving classification accuracy. The overall average accuracy of the optimal model reached 98.99%, providing strong support for the accurate diagnosis of brain tumors.
[0025] 2. Enhanced model generalization: To address the limited and unbalanced nature of medical image data, data augmentation techniques were used to rotate and flip the original images, expanding and balancing the dataset. This significantly increased data diversity and effectively reduced overfitting caused by insufficient data. Furthermore, a 5-fold cross-validation approach was used to evaluate the combined model, verifying its performance from multiple data partitioning perspectives. This further ensured the model's stability under varying data distributions and significantly improved its generalization, enabling it to more reliably classify various types of brain tumor MRI images in actual clinical applications.
[0026] 3. Balancing Computing Resources and Performance: The EfficientNetB0 model utilizes lightweight MBConv modules and depthwise separable convolution techniques to significantly reduce the model's computational overhead. When combined with ResNet50 to form a combined model, this significantly improves model training and inference speed while maintaining high classification accuracy. This enables this method to achieve good performance even with limited computing resources, reducing hardware requirements. This makes it suitable not only for high-performance computing environments but also for use in resource-constrained scenarios, demonstrating its broader applicability and potential for application.
[0027] 4. Optimized Model Structure Design: During the construction of the combined model, the advantages of different models were fully utilized through the rational design of a parallel dual-branch structure and feature fusion. Furthermore, the introduction of L2 regularization and Dropout layers effectively controlled model complexity, prevented the model from overfitting the training data, and enhanced its robustness. This optimized model structure not only improved classification performance but also provided design ideas and technical solutions for similar multi-model combination approaches.
[0028] 5. Comprehensive Performance Evaluation System: During model evaluation, the combined model is comprehensively evaluated from multiple dimensions using multiple metrics, including accuracy, sensitivity, specificity, and AUC. Compared to single-metric evaluation, this more accurately and comprehensively reflects the performance characteristics and strengths and weaknesses of the model, providing a more detailed and reliable basis for model improvement and optimization, and helping to promote the continuous development and improvement of brain tumor MRI image classification technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flow chart of the classification task of the combined model of the present invention;
[0030] Figure 2 This is a flow chart of the architecture of ResNet50 of the present invention;
[0031] Figure 3 This is a flow chart of the architecture of EfficientNetB0 of the present invention;
[0032] Figure 4 This is a flow chart of the architecture of EfficientNetV2B0 of the present invention;
[0033] Figure 5 Sagittal, axial, and coronal images of four different brain tumor types analyzed in the present invention;
[0034] Figure 6 This is a contour-based cropping diagram of the brain region used in the experimental analysis of the present invention;
[0035] Figure 7 Figure 2 is a diagram of the data preprocessing steps in the experimental analysis of the present invention;
[0036] Figure 8 Graph showing the average accuracy (a) and average loss rate (b) of the combined model in the training set and validation set in the experimental analysis of the present invention. DETAILED DESCRIPTION
[0037] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0038] Example 1
[0039] like Figure 1 、 2 As shown in Figure 3, a migration learning method for brain tumor image classification based on multi-model combination is described. The specific technical solution is as follows:
[0040] 1. Image preprocessing
[0041] Utilizing specialized image analysis algorithms, we perform edge contour cropping on acquired brain tumor MRI images. Advanced image segmentation techniques accurately identify brain region boundaries, removing irrelevant background areas from the image and retaining only the primary brain contours. Furthermore, we uniformly scale the processed images to a specific size and format, laying the foundation for subsequent data augmentation and model training, reducing the impact of irrelevant information on classification results and improving the quality and consistency of image data.
[0042] 2. Data Augmentation
[0043] Considering the prevalent problem of small amounts of medical image data and uneven distribution of categories, data augmentation techniques are used to expand and balance the preprocessed images. Specific operations include randomly rotating the images within a certain angle range (e.g., ±15°) and performing horizontal and vertical flipping operations. By combining various data augmentation operations, the diversity of the image data is increased, the dataset size is expanded, and the number of images in each category is more balanced. This effectively alleviates the problem of model overfitting caused by insufficient data and improves the model's generalization ability.
[0044] 3. Build a portfolio model
[0045] 3.1 Model Selection and Architecture Design
[0046] A multi-scale feature fusion strategy was employed to construct a hybrid architecture: ResNet50 + EfficientNetB0. ResNet50 is a classic deep residual network, consisting of multiple residual modules connected in series. Each residual module contains a bottleneck structure consisting of a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer, as well as a batch normalization layer and a Reluctant Unit (ReLU) activation function. This architecture, through residual connections, effectively addresses the vanishing gradient and degradation issues inherent in deep networks, enabling the extraction of rich image features, from low-level edges to high-level semantics. The detailed network structure of ResNet50 is shown in Table 1.
[0047]
[0048]
[0049] Table 1
[0050] EfficientNetB0 primarily utilizes the MBConv module, which draws upon and improves upon the inverted residual architecture of MobileNetV2. Input data sequentially undergoes a 1×1 convolution for dimensionality increase, followed by a K×K depthwise separable convolution, an SE attention module, and a 1×1 convolution for dimensionality reduction. This is followed by a dropout layer, and finally a residual connection to output the feature map. EfficientNetB0 utilizes a compound scaling method to simultaneously and evenly scale the network width, depth, and input image resolution, reducing computational effort while maintaining model performance. This approach is particularly well-suited for scenarios with limited hardware resources. Table 2 shows the detailed network structure of EfficientNetB0.
[0051]
[0052] Table 2
[0053] 3.2 Parallel Dual-branch Structure and Feature Fusion
[0054] ResNet50 and EfficientNetB0 work together in a parallel dual-branch architecture. The preprocessed input image enters both the ResNet50 branch and the EfficientNetB0 branch, where features are extracted. The features extracted by each branch are processed separately through a global average pooling layer, which effectively preserves the image's key features while reducing feature dimensionality. Subsequently, the features extracted by the two branches are concatenated and fused along the channel dimension using a concatenate layer. This allows the model to leverage multi-scale features from both branches for a more comprehensive understanding of the image.
[0055] 3.3 Model optimization layer settings
[0056] After the merging layer after feature fusion, a Dense layer is added, and L2 regularization is introduced. L2 regularization effectively controls model complexity by limiting the size of weights, preventing the model from overfitting to the training data during training. Furthermore, a Dropout layer is added to randomly discard the output of some neurons with a set probability during training, increasing the randomness of the model, further improving its robustness and preventing overfitting. Finally, a SoftMax function is connected to map the fused features to different brain tumor categories, completing the MRI image classification task. The structure and parameters of ResNet50+EfficientNetB0 are shown in Table 3.
[0057]
[0058] Table 3
[0059] 4. Model training and fine-tuning
[0060] Using the dataset obtained after data augmentation, the constructed combined model was trained and fine-tuned based on the principle of transfer learning. First, ResNet50 and EfficientNetB0 were pre-trained on a large-scale general dataset to obtain the pre-trained model parameters. Then, for the brain tumor MRI image classification task, the pre-trained model parameters were used as initial values and the combined model was trained using the augmented dataset.
[0061] During training, select an appropriate optimizer (such as Adamax), set a reasonable learning rate (such as 0.0001-0.001) and number of training rounds (such as 30-80 rounds). Use the backpropagation algorithm to continuously update model parameters and minimize the loss function (such as the cross-entropy loss function), gradually adapting the model to the brain tumor image classification task, extracting more discriminative features, and optimizing model performance.
[0062] 5. Model Evaluation
[0063] The trained and fine-tuned combined model was tested using 5-fold cross-validation on the Kaggle dataset. The dataset was divided into five equal parts, with four parts used as training sets and one part used as testing sets. Training and testing were repeated five times. Accuracy, sensitivity, specificity, AUC, and other evaluation metrics were calculated for each test, and the average was taken as the final evaluation result.
[0064] The classification performance of this combined model was also compared with that of three single CNN models (ResNet50, EfficientNetB0, and EfficientNetV2B0), as well as existing popular brain tumor classification methods. The model's strengths and weaknesses were comprehensively evaluated across multiple dimensions, including classification accuracy, computational efficiency, and generalization, ensuring the reliability and stability of the results and validating the effectiveness and superiority of this method in brain tumor MRI image classification.
[0065] Example 2
[0066] like Figure 1 、 2 As shown in Figure 4, a migration learning method for brain tumor image classification based on multi-model combination is described. The specific technical solution is as follows:
[0067] 1. Image preprocessing
[0068] Utilizing specialized image analysis algorithms, we perform edge contour cropping on acquired brain tumor MRI images. Advanced image segmentation techniques accurately identify brain region boundaries, removing irrelevant background areas from the image and retaining only the primary brain contours. Furthermore, we uniformly scale the processed images to a specific size and format, reducing the impact of irrelevant information on classification results and improving the quality and consistency of the image data, laying the foundation for subsequent data enhancement and model training.
[0069] 2. Data Augmentation
[0070] Given the limited amount of medical image data and the uneven distribution of categories, data augmentation techniques are used to augment and balance the preprocessed images. Specific operations include random rotation of the images within a ±15° range and horizontal and vertical flipping. By combining various data augmentation operations, we increase the diversity of the image data, expand the dataset, and achieve a more balanced number of images across categories. This effectively mitigates model overfitting caused by insufficient data and improves the model's generalization capabilities.
[0071] 3. Build a portfolio model
[0072] 3.1 Model Selection and Architecture Design
[0073] A multi-scale feature fusion strategy is adopted to construct a hybrid architecture ResNet50+EfficientNetV2B0.
[0074] ResNet50 is a classic deep residual network composed of multiple residual modules connected in series. Each residual module contains a bottleneck structure consisting of a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer, as well as a batch normalization layer and a ReLU activation function. Residual connections effectively address the vanishing gradient and degradation issues of deep networks, enabling the extraction of rich image features, from low-level edges to high-level semantics.
[0075] EfficientNetV2B0 is the base model in the EfficientNetV2 series. It replaces the first three MBConv modules with Fused-MBConv modules. In this Fused-MBConv module, the original 1×1 convolution and depthwise separable convolution are replaced with 3×3 convolutions. When the channel expansion rate of the input image is not equal to 1, an additional 1×1 convolution layer is added. The EfficientNetV2B0 network architecture consists of eight stages. Stage 0 uses a 3×3 convolution layer to generate basic features. Stages 1 through 3 utilize the Fused-MBConv module to improve computational and feature extraction capabilities. Stages 4 through 6 utilize the MBConv module to mine more complex features. Stage 7 completes the final classification task with a 1×1 convolution layer, a global average pooling layer, and a fully connected layer. This model utilizes the Fused-MBConv module and a progressive learning approach to reduce computational complexity, improve training and inference speed, and maintain high performance, making it particularly suitable for resource-constrained environments. The detailed network structure of EfficientNetV2B0 is shown in Table 4.
[0076]
[0077] Table 4
[0078] 3.2 Parallel Dual-branch Structure and Feature Fusion
[0079] ResNet50 and EfficientNetV2B0 work together in a parallel dual-branch architecture. The preprocessed input image is fed into both the ResNet50 and EfficientNetV2B0 branches for feature extraction. The features extracted by each branch are processed separately through a global average pooling layer, preserving key image features while reducing feature dimensionality. Subsequently, the features extracted by the two branches are concatenated and fused channel-wise using a concatenate layer. This allows the model to leverage multi-scale features from both branches for a more comprehensive understanding of the image.
[0080] 3.3 Model optimization layer settings
[0081] After the merging layer after feature fusion, a fully connected layer (Dense layer) is added, and L2 regularization is introduced. L2 regularization controls model complexity by limiting the size of weights, preventing the model from overfitting the training data during training. Furthermore, a Dropout layer is added to randomly discard the outputs of some neurons with a set probability during training, increasing the model's randomness, further improving its robustness, and avoiding overfitting. Finally, a SoftMax function is connected to map the fused features to different brain tumor categories, completing the MRI image classification task. The combined model architecture has a total of 30,360,276 parameters, of which 30,246,548 are optimizable parameters used for weight updates in the feature extraction and classification modules; 113,728 are fixed parameters, which come from specific frozen layers of the pretrained model. The structure and parameters of ResNet50+EfficientNetV2B0 are shown in Table 5.
[0082]
[0083]
[0084] Table 5
[0085] 4. Model training and fine-tuning
[0086] Using the data augmentation dataset, we trained and fine-tuned the constructed ResNet50 + EfficientNetV2B0 combined model based on transfer learning. First, we pre-trained ResNet50 and EfficientNetV2B0 on a large-scale general dataset to obtain the pre-trained model parameters. Then, for the brain tumor MRI image classification task, we used the augmented dataset, using the pre-trained model parameters as initial values to train the combined model.
[0087] During training, the Adamax optimizer was selected, with a learning rate of 0.0001-0.001 and 30-80 training rounds. The backpropagation algorithm continuously updated model parameters to minimize the cross-entropy loss function, gradually adapting the model to the brain tumor image classification task, extracting more discriminative features, and optimizing model performance.
[0088] 5. Model Evaluation
[0089] The trained and fine-tuned combined model was tested using 5-fold cross-validation on the Kaggle dataset. The dataset was divided into five equal parts, with four parts used as training sets and one part used as testing sets. Training and testing were repeated five times. Accuracy, sensitivity, specificity, AUC, and other evaluation metrics were calculated for each test, and the average was taken as the final evaluation result.
[0090] The classification performance of this combined model was compared with that of three single CNN models: ResNet50, EfficientNetB0, and EfficientNetV2B0; a combined ResNet50 + EfficientNetB0 model; and existing popular brain tumor classification methods. The model's strengths and weaknesses were comprehensively evaluated across multiple dimensions, including classification accuracy, computational efficiency, and generalization, ensuring the reliability and stability of the results and validating the effectiveness and superiority of this method in brain tumor MRI image classification.
[0091] Example 3
[0092] like Figure 1 、 3 As shown in Figure 4, a migration learning method for brain tumor image classification based on multi-model combination is described. The specific technical solution is as follows:
[0093] 1. Image preprocessing
[0094] Using specialized image analysis algorithms, we perform edge contour cropping on acquired brain tumor MRI images. Leveraging advanced image segmentation techniques, we accurately identify brain region boundaries, remove irrelevant background areas, and retain only the primary brain contours. Furthermore, we uniformly scale the processed images to a specific size and format, minimizing the impact of irrelevant information on classification results. This improves the quality and consistency of the image data, laying a solid foundation for subsequent data enhancement and model training.
[0095] 2. Data Augmentation
[0096] Given the limited amount of medical image data and the uneven distribution of categories, data augmentation techniques are used to augment and balance the preprocessed images. Specific operations include random rotation of the images within a ±15° range and horizontal and vertical flipping. By combining various data augmentation operations, we increase the diversity of the image data, expand the dataset, and achieve a more balanced number of images across categories. This effectively mitigates model overfitting caused by insufficient data and improves the model's generalization capabilities.
[0097] 3. Build a portfolio model
[0098] 3.1 Model Selection and Architecture Design
[0099] A multi-scale feature fusion strategy is adopted to construct the hybrid architecture EfficientNetB0+EfficientNetV2B0.
[0100] EfficientNetB0 consists of nine stages, of which the second to eighth stages are mainly composed of 16 stacked MBConv modules. The MBConv module draws on and further improves the inverted residual structure of MobileNetV2. The input data passes through a 1×1 convolution (for dimensionality increase), a K×K depthwise convolution, an SE module, and a 1×1 convolution (for dimensionality reduction). After that, the main branch of the module is randomly reduced through a Dropout layer to reduce the network depth. Finally, a residual connection is performed to output the feature map. The model uses advanced technologies such as depthwise separable convolution, compound scaling strategy, Swish activation function, and SE (Squeeze-and-Excitation) module to effectively improve the network's operating efficiency and performance.
[0101] EfficientNetV2B0 is the base model in the EfficientNetV2 series. It replaces the first three MBConv modules with Fused-MBConv modules. In this Fused-MBConv module, the original 1×1 convolution and depthwise separable convolution are replaced with 3×3 convolutions. When the channel expansion rate of the input image is not equal to 1, an additional 1×1 convolution layer is added. The network architecture of EfficientNetV2B0 consists of eight stages. Stage 0 uses a 3×3 convolution layer to generate basic features. Stages 1 through 3 utilize the Fused-MBConv module to improve computational and feature extraction capabilities. Stages 4 through 6 utilize the MBConv module to mine more complex features. Stage 7 completes the final classification task with a 1×1 convolution layer, a global average pooling layer, and a fully connected layer. This model utilizes the Fused-MBConv module and a progressive learning approach to reduce computational effort and improve training and inference speed while maintaining high model performance, making it particularly suitable for resource-constrained environments.
[0102] 3.2 Parallel Dual-branch Structure and Feature Fusion
[0103] EfficientNetB0 and EfficientNetV2B0 work together in a parallel dual-branch architecture. A preprocessed 224×224×3 image is fed into both the EfficientNetB0 and EfficientNetV2B0 branches for feature extraction. After feature extraction, both branches perform global average pooling on the resulting 7×7 feature maps. A concatenate layer then concatenates and fuses the features from both branches along the channel dimension. This allows the model to leverage multi-scale features from both branches, achieving a more comprehensive understanding of the image.
[0104] 3.3 Model optimization layer settings
[0105] After the merging layer after feature fusion, a fully connected layer (Dense layer) is added, and L2 regularization is introduced. L2 regularization controls model complexity by limiting the size of weights, preventing the model from overfitting to the training data during training. Furthermore, a Dropout layer is added to randomly discard the outputs of some neurons with a set probability during training, increasing the model's randomness, further improving its robustness, and preventing overfitting. Finally, a SoftMax function is connected to map the fused features to different brain tumor categories, completing the MRI image classification task. This combined model has a total of 10,625,527 parameters, of which 10,522,896 are trainable parameters and 102,631 are non-trainable parameters. The structure and parameters of EfficientNetB0 + EfficientNetV2B0 are shown in Table 6.
[0106]
[0107] Table 6
[0108] 4. Model training and fine-tuning
[0109] Using the data-augmented dataset, we trained and fine-tuned the constructed EfficientNetB0 + EfficientNetV2B0 combined model based on transfer learning. First, we pre-trained EfficientNetB0 and EfficientNetV2B0 on a large-scale, general-purpose dataset to obtain the pre-trained model parameters. Then, for the brain tumor MRI image classification task, we used the augmented dataset, using the pre-trained model parameters as initial values to train the combined model.
[0110] During training, the Adamax optimizer was selected, with a learning rate of 0.0001-0.001 and 30-80 training rounds. The backpropagation algorithm continuously updated model parameters to minimize the cross-entropy loss function, gradually adapting the model to the brain tumor image classification task, extracting more discriminative features, and optimizing model performance.
[0111] 5. Model Evaluation
[0112] The trained and fine-tuned combined model was tested using 5-fold cross-validation on the Kaggle dataset. The dataset was divided into five equal parts, with four parts used as training sets and one part used as testing sets. Training and testing were repeated five times. Accuracy, sensitivity, specificity, AUC, and other evaluation metrics were calculated for each test, and the average was taken as the final evaluation result.
[0113] The classification performance of this combined model was compared with that of three single CNN models: ResNet50, EfficientNetB0, and EfficientNetV2B0; as well as the combined models of ResNet50+EfficientNetB0 and ResNet50+EfficientNetV2B0; and existing popular brain tumor classification methods. The strengths and weaknesses of the models were comprehensively evaluated across multiple dimensions, including classification accuracy, computational efficiency, and generalization, to ensure the reliability and stability of the results and verify the effectiveness and superiority of this method in brain tumor MRI image classification.
[0114] Experimental analysis
[0115] 1. Dataset and Preprocessing Methods
[0116] 1.1 Dataset Description
[0117] The data used in this paper comes from the "brain MR images" dataset on Kaggle. This dataset contains four types of T1-weighted contrast-enhanced images, namely glioma images (926 images), meningioma images (937 images), pituitary tumor images (901 images), and tumor-free images (500 images). The number of categories in the dataset is shown in Table 7. In addition, Figure 5 The dataset shows three image planes from four different categories: sagittal, axial, and coronal. The sagittal plane, which runs along the front-to-back axis of the body, facilitates visualization of the brain's anterior-posterior structures. The axial plane, which runs horizontally, is useful for observing the brain's superior-inferior structures. The coronal plane, which runs along the left-to-right axis of the body, facilitates visualization of the brain's left-to-right structures. Three-dimensional image information provides more comprehensive anatomical information, which is crucial for the accurate diagnosis and analysis of brain diseases.
[0118]
[0119] Table 7
[0120] 1.2 Data Partitioning
[0121] To ensure the generalization and reliability of the model, this paper uses a five-fold cross-validation method to partition the dataset. First, the entire data sample is randomly shuffled to eliminate any possible order bias in the data, thereby ensuring the representativeness of each subset. Then, using five-fold cross-validation, the shuffled dataset is divided into five subsets. In each round of five-fold cross-validation, one of the subsets (20%) is selected as the test set, and the remaining four subsets (80%) are combined and used for training.
[0122] During training, to further optimize the model's hyperparameters and prevent overfitting, 20% of the training data is set aside as a validation set for hyperparameter tuning and training monitoring. This allows for timely detection and correction of potential overfitting issues during training, thereby improving the model's generalization capabilities.
[0123] This partitioning approach ensures that the model is fully trained while ensuring that every sample participates in the test. It also effectively prevents overfitting through an independent validation set, making the evaluation results more statistically significant and reliable. The final model performance is the average of the five test rounds, providing a more comprehensive assessment of the model's performance on unseen data.
[0124] 1.3 Data Preprocessing
[0125] In order to reduce the waste of computing resources, the MRI image is cropped to retain only the main contour area of the brain, such as Figure 6 As shown in the figure, the image is first grayscaled. Subsequently, threshold segmentation is applied, and a series of erosion and dilation operations are performed to further reduce image noise and enhance contour clarity. Next, the maximum contour is identified based on the threshold map, and the contour's extreme points (top, bottom, left, and right) are determined. Finally, the image is cropped at these endpoints, effectively removing background interference and noise. After image cropping, the input original image is scaled to 224×224 pixels.
[0126] Because the number of images in the training set is uneven across different categories, and the total number of samples is insufficient to effectively train a deep CNN model, we use data augmentation techniques such as random image rotation to generate additional slices for each category. In the Kaggle dataset, for the "NoTumor" category, which has the fewest samples, we use data augmentation techniques such as random image rotation to generate additional slices, expanding the size of the data to four times the original size. For the other three tumor categories, we use the same augmentation method to expand the size of the data to twice the original size. The specific data augmentation operations are shown in Table 8.
[0127]
[0128] Table 8
[0129] The data preprocessing process is as follows Figure 7 shown.
[0130] After completing the above operations, the input image is normalized to ensure the stability and efficiency of model training. This paper adopts the Min-Max normalization method to rescale the input data to the range of (0, 1). The formula is as follows:
[0131]
[0132] Among them, x i represents the specific value of a numerical feature of the i-th sample, and x min and x max They represent the minimum and maximum values of the feature in all samples respectively.
[0133] 2. Experimental environment and hyperparameter configuration
[0134] This article will outline the hardware and software environment required for training and evaluating the model. The selection of these hardware and software resources is crucial for ensuring efficient model training and accurate evaluation. Next, we will explore in detail the selection criteria and value ranges of the hyperparameters used in the model training process, as well as their impact on the model training process and final performance. Finally, we will describe the performance metrics used in the model evaluation phase. These metrics reflect the model's performance in various aspects, providing a comprehensive and accurate basis for model performance evaluation.
[0135] 2.1 Experimental Configuration
[0136] The experimental setup for this paper includes a high-performance hardware and software environment. For hardware, we used an Intel Core i7-12650H processor, 16GB of RAM, and a GeForce RTX 3050 graphics card. For software, we used Python 3.10 and the TensorFlow 2.17.1 deep learning framework.
[0137] 2.2 Experimental performance evaluation indicators
[0138] In medical diagnostic systems, choosing the right evaluation metric is crucial for measuring model performance. This is crucial not only for obtaining correct results but also for avoiding misleading ones. Therefore, selecting a performance metric is a crucial step in objectively evaluating model success.
[0139] The present invention uses precision, sensitivity, specificity, accuracy, F1 score and AUC value to evaluate the classification performance of the model. Precision measures the reliability of the prediction when diagnosing a disease, that is, how many cases predicted to be diseased actually have the disease. Sensitivity describes the proportion of samples that are correctly predicted to be positive among all samples that are actually positive. Specificity measures the ability of the model to correctly identify negative samples, that is, the proportion of samples that are correctly predicted to be negative among all samples that are actually negative. Accuracy indicates the ability of the system to make a correct diagnosis, that is, the ratio of correct and incorrect diagnoses of the system, and is the overall classification accuracy of the proposed method in terms of TP and TN. The F1 score is an indicator that measures classification performance from two aspects: sensitivity and precision. The specific evaluation indicators are as follows:
[0140] Precision (Pr)
[0141] Precision refers to the ratio of the number of samples correctly predicted by the model to the total number of samples predicted by the model to be of that category. It reflects the accuracy of the model in predicting a certain category. The calculation formula is:
[0142]
[0143] Sensitivity (Se)
[0144] Sensitivity describes the proportion of positive samples correctly identified by the model to the total number of actual positive samples. It reflects the model's ability to identify a specific category. The calculation formula is:
[0145]
[0146] Specificity (Sp)
[0147] Specificity refers to the ratio of the number of samples correctly predicted by the model as not belonging to a certain category to the total number of samples that are not actually belonging to that category. It reflects the accuracy of the model in excluding a certain category. The calculation formula is:
[0148]
[0149] Accuracy (Acc)
[0150] Accuracy refers to the ratio of the number of samples correctly predicted by the model to the total number of samples. It reflects the overall accuracy of the model across all categories. The calculation formula is:
[0151]
[0152] F1 score (F1)
[0153] The F1 score is the harmonic mean of precision and recall (sensitivity), which takes into account the balance between precision and recall. The calculation formula is:
[0154]
[0155] Where TP = True Positives, which represents the number of positive examples correctly identified by the model. TN = True Negatives, which represents the number of negative examples correctly identified by the model. FP = False Positives, which represents the number of negative examples incorrectly identified as positive by the model. FN = False Negatives, which represents the number of positive examples incorrectly identified as negative by the model.
[0156] 2.3 Parameter settings
[0157] Based on the characteristics of deep neural networks and the requirements of the target task, this paper systematically configures and optimizes the model's hyperparameters. Table 9 summarizes all hyperparameters used in this study and their corresponding values.
[0158]
[0159]
[0160] Table 9
[0161] The specific experimental settings are as follows:
[0162] (1) Training settings
[0163] a. Epochs: Iterations refer to the number of times the model is fully trained on the entire training dataset. In each iteration, the model traverses the entire training set once and updates the parameters through backpropagation to reduce prediction errors. An appropriate number of iterations helps the model gradually learn the patterns and features in the data, thereby better fitting the training data. However, too many iterations may cause the model to overfit, that is, it performs well on the training data, but the performance degrades on new data; while too few iterations may cause underfitting, that is, the model cannot fully learn the complex relationships in the data. Therefore, choosing an appropriate number of iterations requires finding a balance between underfitting and overfitting. The number of iterations in this article is set to 30. This setting ensures that the model has enough time to learn the features in the data while avoiding overtraining.
[0164] b. Batch Size: Batch size refers to the number of data samples used each time a model parameter is updated during training. It is a critical hyperparameter in deep learning training, directly impacting training efficiency and model performance. Larger batch sizes can improve computational efficiency during training, but they can also cause the model to become trapped in a local minimum during training. Conversely, smaller batch sizes, while improving model generalization, can result in lower computational efficiency and potentially unstable training. Therefore, selecting an appropriate batch size balances training efficiency and model performance to achieve optimal training results. In our experiments, to ensure efficient resource utilization for deep learning model training without exceeding hardware limitations, we carefully adjusted the batch size to accommodate the performance constraints of the experimental equipment. After testing, we set the batch size to 16, the maximum batch size supported by the experimental equipment.
[0165] (2) Optimization algorithm
[0166] The Adamax optimizer was chosen for its efficiency and stability in deep network training, making it particularly suitable for processing complex neural network structures. The Adamax optimizer is a variant of the Adam optimization algorithm that adjusts the learning rate based on the infinite norm instead of the L2 norm in Adam.
[0167] (3) Learning rate setting
[0168] The learning rate is a key hyperparameter in deep learning training, which determines the step size of the model parameters updated in each iteration. An appropriate learning rate can accelerate the convergence of the model while avoiding instability during training. The initial learning rate of the present invention is set to 0.001, which is a commonly used default value and is suitable for most deep learning tasks. However, a fixed learning rate may perform poorly at different stages of training: in the early stages of training, a larger learning rate helps to converge quickly, but in the later stages of training, a smaller learning rate helps to adjust the model parameters more finely and avoid falling into local optimality.
[0169] To dynamically adjust the learning rate during training, this study introduced the ReduceLROnPlateau learning rate scheduler. This scheduler dynamically adjusts the learning rate during training: if validation set accuracy does not improve for five consecutive epochs, the learning rate is reduced to 0.1 times the original value. This dynamic adjustment strategy allows the model to converge quickly using a larger learning rate at the beginning of training, and then, when validation set performance stagnates, to more carefully optimize model parameters by reducing the learning rate.
[0170] (4) Model architecture and regularization
[0171] After the global average pooling layer, a concatenation layer is used to concatenate the features extracted by the two models. A fully connected layer (Dense Layer) with 256 neurons is then used, and L2 regularization is introduced to control model complexity. To further prevent overfitting, a dropout layer is also used with a dropout rate of 0.5.
[0172] (5) Output layer
[0173] SoftMax is often used in the output layer of multi-classification problems. Its output is a probability distribution, that is, the sum of all elements is equal to 1, and each element represents the probability of a category. The calculation formula is:
[0174]
[0175] Where S(x i ) represents the predicted probability of the i-th category. i Represents the original output of the i-th category, and n is the number of categories.
[0176] (6) Loss function
[0177] The loss function used is categorical cross entropy. Categorical cross entropy combines the SoftMax activation function with cross entropy loss and is suitable for multi-classification problems. Its purpose is to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label. Only the positive class contributes to the cross entropy loss. The calculation formula is as follows:
[0178]
[0179] Among them, y i is the value of the true label in the i-th category in the one-hot encoding. For the positive class (ie, the correct category), y i =1, for other categories, y i = 0. S(x i ) is the SoftMax function, which is the probability of the i-th category predicted by the model.
[0180] 3. Experimental results and performance analysis
[0181] To comprehensively evaluate the strengths and weaknesses of the combined models, we conducted comparative experiments using 5-fold cross-validation to compare the combined models with individual models. Using the same dataset and parameters, we compared ResNet50+EfficientNetB0, ResNet50+EfficientNetV2B0, and EfficientNetB0+EfficientNetV2B0 with ResNet50, EfficientNetB0, and EfficientNetV2B0. The results of these comparative experiments are shown in Table 10. The experimental results show that among the three combined models, the ResNet50+EfficientNetB0 combined model exhibits the best classification performance, achieving an overall accuracy of 98.99%, with a classification accuracy of 99.85% in the fourth fold. EfficientNetB0+EfficientNetV2B0 ranked second with an overall accuracy of 98.84%, while ResNet50+EfficientNetV2B0 achieved the lowest overall accuracy of 98.25%. From the perspective of model structure, more parameters does not mean better classification results. Taking ResNet50+EfficientNetV2B0 as an example, this model has the most parameters and the highest model complexity, which can easily lead to model overfitting, thereby affecting classification accuracy. However, too few parameters will make it difficult for the model to learn the deep features of the image and fail to pay attention to subtle lesions. This can be confirmed by the experimental results of single models. The classification accuracy of single models is generally lower than that of combined models. The overall accuracy of ResNet50, EfficientNetB0, and EfficientNetV2B0 are 96.57%, 96.94%, and 95.53%, respectively.
[0182]
[0183] Table 10
[0184] Figure 8The figure shows the average accuracy and loss value curves of the method of the present invention in the training set and the validation set. The x-axis represents the round, and the y-axis represents the degree of improvement. The training curves drawn in the figure show the training effect of the model. The models all showed good learning ability during the training process. Their training accuracy rates rose rapidly and stabilized, approaching 1.0. The validation accuracy rate also entered a stable stage after rising and remained at a high level. As the number of training rounds (epochs) increases, it can be seen that the loss value of the model also tends to decrease synchronously. These learning curves show that the models are able to learn and process the given input data well in each training stage, showing that these models have good generalization capabilities on both the training set and the validation set.
[0185] In order to comprehensively evaluate the performance of the model, the present invention uses five-fold cross-validation to calculate multiple key indicators for each category, including precision, sensitivity, specificity, F1-score, and average accuracy. Table 11 summarizes the average results of the five-fold cross-validation to more clearly demonstrate the performance of each model.
[0186]
[0187]
[0188] Table 11
[0189] The ResNet50+EfficientNetB0 combined model demonstrated the most balanced performance across all categories. In particular, in the detection of pituitary tumors, its sensitivity reached 100%, meaning the model was able to identify all pituitary tumor samples. Furthermore, the AUC values for the three combined models were 99.90%, 99.83%, and 99.92%, respectively, demonstrating that these models were able to effectively distinguish between different categories of brain tumor images. The AUC value is an important metric for evaluating the performance of a classification model, reflecting the overall performance of the model under all possible classification thresholds. A high AUC value indicates that the model has excellent generalization capabilities and can accurately distinguish between samples of different categories.
[0190] In summary, from model design, data processing, experimental setup, and experimental validation, we systematically demonstrated the performance and advantages of three combined models: ResNet50+EfficientNetB0, ResNet50+EfficientNetV2B0, and EfficientNetB0+EfficientNetV2B0. Experiments were also conducted to compare these combined models with the individual models: ResNet50, EfficientNetB0, and EfficientNetV2B0. The experimental results showed that the optimal combined model, ResNet50+EfficientNetB0, achieved an overall accuracy of 98.99%. Furthermore, the combined model demonstrated superior classification performance for brain tumor MRI images compared to the individual models. We then compared the classification performance of each architecture for glioma, meningioma, pituitary tumor, and non-tumor categories.
[0191] These results show that the combined model of feature fusion surpasses the accuracy and reliability of single models. This model combines the advantages of different architectures, such as the deep feature extraction of ResNet50 and the computational efficiency of EfficientNetB0, thereby improving the efficiency and accuracy of disease diagnosis.
[0192] The present invention is provided as an example, not as a limitation of the embodiments. Those skilled in the art will appreciate that other variations or modifications may be made based on the above description. It is not necessary and impossible to enumerate all embodiments here, and obvious variations or modifications derived therefrom remain within the scope of protection of the present invention.
Claims
1. A multi-model combined transfer learning method for brain tumor image classification, characterized by: The following steps are involved: Image preprocessing: perform edge contour cropping and scaling on brain tumor MRI images to retain the main contour area of the brain; Data enhancement: Using data enhancement technology, we rotate and flip the pre-processed images, expand and balance the original brain tumor MRI images, and obtain an enhanced dataset. Building a combined model: A multi-scale feature fusion strategy is used to construct a hybrid architecture, ResNet50+EfficientNetB0. ResNet50 and EfficientNetB0 work together in a parallel dual-branch structure. The input image after image preprocessing is simultaneously fed into both branches for feature extraction. After being processed by the global average pooling layer, the extracted features are concatenated through the Concatenate layer. A Dense layer is added after the merging layer, and L2 regularization is introduced. A Dropout layer is added, and finally a SoftMax function is used for multi-classification processing. Model training and fine-tuning: Using the enhanced dataset obtained in the data augmentation step, based on the idea of transfer learning and the pre-trained model parameters, the constructed combined model is trained and fine-tuned for the brain tumor MRI image classification task; Finally, the model is evaluated: 5-fold cross-validation is used on the Kaggle dataset to test the combined model after model training and fine-tuning, and the classification performance is compared and analyzed with other models and existing popular methods.
2. The method for brain tumor image classification based on multi-model combination according to claim 1, characterized in that: The ResNet50 model is composed of multiple residual modules connected in series. Each residual module contains a "bottleneck structure" consisting of a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer, as well as a batch normalization layer and a ReLU activation function.
3. The method for brain tumor image classification based on multi-model combination according to claim 1, characterized in that: The EfficientNetB0 model consists of 9 stages, of which the second to eighth stages are formed by stacking 16 MBConv modules, which sequentially include 1×1 convolution for dimensionality increase, K×K depthwise separable convolution, SE attention module, 1×1 convolution for dimensionality reduction, Dropout layer and residual connection.
4. The method for brain tumor image classification transfer learning based on multi-model combination according to claim 1, characterized in that: In the parallel dual-branch structure, the features extracted by the ResNet50 branch and the EfficientNetB0 branch are processed by the global average pooling layer and then spliced and fused according to the channel dimension.
5. The method for brain tumor image classification transfer learning based on multi-model combination according to claim 1, characterized in that: During the model training and fine-tuning process, the Adamax optimizer was used, the learning rate was set to 0.0001-0.001, and the training rounds were set to 30-80 rounds.
6. The method for brain tumor image classification transfer learning based on multi-model combination according to claim 1, characterized in that: In the model evaluation step, in addition to accuracy, sensitivity, specificity, and AUC value indicators are also used to evaluate the performance of the combined model.
7. The method for brain tumor image classification transfer learning based on multi-model combination according to claim 1, characterized in that: When building a combined model, you can also choose the ResNet50+EfficientNetV2B0 or EfficientNetB0+EfficientNetV2B0 architecture to achieve brain tumor MRI image classification through the feature combination of the corresponding model branches.
Citation Information
Cited By
Tumor pathological image classification method and system based on artificial intelligence
CN121170436A