Image classification method based on improved Inception v4 model
By embedding the CBAM attention mechanism and FRN layer in the Inception v4 model and using the cross entropy loss function, the feature extraction and training methods of the model are optimized, and the problem of insufficient feature extraction capabilities and risk of overfitting is solved, and higher image classification accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510224971.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The existing Inception v4 model has problems with insufficient feature extraction capabilities and risk of overfitting in image classification tasks, resulting in low classification accuracy and efficiency.
By embedding the CBAM attention mechanism and FRN layer in the Inception v4 model and using the cross entropy loss function, the feature extraction and training methods of the model are optimized to improve the feature fusion ability and robustness of the model.
Without significantly increasing the number of parameters, the feature extraction quality and classification accuracy of the model are improved, the risk of overfitting is reduced, and the efficiency of image classification is improved.
Smart Images

Figure CN120071008A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image classification, and specifically relates to an image classification method based on an improved Inception v4 model Background Art
[0002] Image classification is a core task in the field of computer vision, aiming to construct a system or model that can automatically identify and classify the content of images. It promotes the development of computer vision technologies (such as object detection, semantic segmentation, image generation, etc.), and has practical application significance in improving production efficiency, enhancing the quality of life, and improving the efficiency of information retrieval
[0003] Currently, with the continuous development of computer technology, the main methods in the field of image classification are divided into two types. One is the method based on machine learning. First, artificial feature extraction algorithms such as Scale-Invariant Feature Transform (SIFT), Histogram of Oriented Gradients (HOG), etc. are designed to extract image features, and then the extracted features are input into traditional machine learning classifiers such as Support Vector Machine (SVM), Decision Tree (RT), etc. for classification. The other is the method based on deep learning. The most common one is the method based on CNN, which is composed of components such as convolutional layers, pooling layers, and fully connected layers. Due to its advantage of automatically extracting image features, it has been widely used in the field of image classification. Network structures such as AlexNet, VGGNet, ResNet, Inception, etc. have achieved excellent results in various image classification tasks. The Inception series of networks can capture multi-scale information by stacking multiple parallel convolutional branch modules to increase the network width. Among them, the Inception v4 model was proposed in 2016. This network structure has classification accuracy comparable to that of Inception-Resnet-V2 without using the residual structure, and has fewer parameters and a more lightweight network Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] To solve the problems in the background art, the purpose of the present invention is to provide an image classification method based on an improved Inception v4 model, which can improve the original Inception v4 accuracy, improve the backbone network architecture of Inception v4, optimize the training method, enhance the image classification efficiency, and improve the model classification accuracy
[0006] (2) Technical Solutions
[0007] The present invention specifically adopts the following technical solutions to achieve the above purpose
[0008] An image classification method based on an improved Inception v4 model, comprising the following steps:
[0009] Step 1: Construct a data set, which includes a training set, a validation set, and a test set.
[0010] Step 2: Improve the Inception v4 model. The specific improvement measures are as follows:
[0011] (1) To address the problem of insufficient feature extraction ability in the Inception-v4 algorithm model, the CBAM attention mechanism is used to enhance the feature extraction ability of the Inception module. Compared with traditional attention mechanisms, CBAM enables the network to selectively enhance the features beneficial to the task in both the spatial and channel dimensions, improving the quality of feature extraction, optimizing the feature space, and making the feature fusion ability of the Inception-v4 algorithm model stronger without significantly increasing the number of parameters.
[0012] (2) Embed the FRN layer in the Reduction module of the Inception v4 model. Different from BN and LN, it does not rely on batch statistical information and can better capture the feature differences between different filters in the convolutional layer. This method improves the convergence speed of the network model, enhances interpretability and robustness, reduces the risk of overfitting, and solves the problem of image category imbalance.
[0013] (3) Determine the cross-entropy loss function. The loss function is used to evaluate the deviation between the predicted value and the true value in the convolutional neural network. The cross-entropy function is a convex function with respect to the predicted probability q(x), which has good properties during the optimization process, ensuring that the gradient-based optimization algorithm can find the global optimal solution or at least a good local optimal solution. In actual calculations, the calculation of cross-entropy is relatively simple, only involving logarithmic operations and multiplication operations, which are easy to implement on a computer and have high calculation efficiency.
[0014] Step 3: Based on the completed model construction, input the training set into the improved network model for training, adjust the parameters to optimize the model, and output the image classification model after training.
[0015] Step 4: Compare the improved model with the original Inception v4 model. After using the same data for training, use accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1_score) to compare and evaluate the models respectively.
[0016] Further, in step 1 of the present invention, prepare the data annotation format required by the Inception network, and divide the entire dataset of images and labels into a training set, a validation set, and a test set according to 8:1:1.
[0017] Further, in (1) of step 2 of the present invention, embed the CBAM attention mechanism in the Inception module of Inception v4. This mechanism processes the input feature map in the channel dimension, calculates the channel attention weights. Multiply the input feature map by the channel attention weights to obtain the feature map weighted by channel attention. Input the feature map weighted by channel attention into the spatial attention module for processing in the spatial dimension, calculate the spatial attention weights. Multiply the feature map weighted by channel attention by the spatial attention weights to obtain the final feature map processed by CBAM. Finally, use the feature map processed by CBAM as the input to the next layer of the convolutional neural network to continue subsequent feature extraction and processing.
[0018] Further, in (2) of step 2 of the present invention, embed the FRN layer in the Reduction module of the Inception v4 model. The FRN layer normalizes the feature map output by the convolution, eliminates the excessive influence of certain features on the model caused by the large difference in feature scales, accelerates the training speed of the model, improves the convergence efficiency, and can also enhance the stability and generalization ability of the model.
[0019] Further, in (3) of step 2 of the present invention, use Cross-Entropy Loss as the loss function. The expression of the Cross-Entropy Loss function for the discrete case is:
[0020] H(p,q) = -∑p(x)log q(x)
[0021] The expression of the Cross-Entropy Loss function for the continuous case is:
[0022] H(p,q) = -∫p(x)log q(x)dx
[0023] Where p(x) is the true probability, q(x) is the predicted probability of the model, x represents different classes. If it is a binary classification problem, assume the true label of the positive class is 1, the probability is p, the true label of the negative class is 0, the probability is 1 - p, and the probability that the model predicts the positive class is The probability of the negative class is
[0024] Further, in step 3 of the present invention, obtain the improved image classification model, train the model and adjust the parameters to optimize the model, and output the image classification model after training is completed.
[0025] Further, in step 4 of the present invention, in order to evaluate the advantages of the improved model, the improved model is compared with the original Inception v4 model. After training with the same dataset, the accuracy, precision, recall, and F1 score are used to compare and evaluate the model respectively.
[0026] The evaluation index formulas are as follows:
[0027]
[0028] Among them, TP represents the number of samples correctly predicted as positive examples, TN represents the number of samples correctly predicted as negative examples by the model, FP represents the number of samples wrongly predicted as positive examples by the model, and FN represents the number of samples wrongly predicted as negative examples by the model. Description of the Drawings
[0029] Figure 1 is the flowchart of the present invention;
[0030] Figure 2 is the schematic diagram of the CBAM attention mechanism
[0031] Figure 3 is the schematic diagram of the FRN layer
[0032] Figure 4 is the network structure diagram of the improved Inception v4 Detailed Embodiments
[0033] The present invention will be described in detail below in conjunction with the detailed embodiments.
[0034] As Figure 1 shown, the present invention provides an image classification method based on an improved Inception v4 model, including the following steps:
[0035] Step 1: Construct a dataset, prepare the data annotation format required by the Inception v4 network, and divide the dataset images and labels into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0036] Step 2: Improve the Inception v4 model. The specific improvement measures are as follows:
[0037] (1) Embed the CBAM attention mechanism in the Inception module of Inception v4. Its schematic diagram is as Figure 2 shown.
[0038] The CBAM attention mechanism integrates two attention modules: channel and spatial. The output feature map of a certain layer in the convolutional neural network serves as the input to CBAM, which has a certain number of channels and a spatial dimension. In the channel attention module, the input feature map is processed in the channel dimension to calculate the channel attention weights. The input feature map is multiplied by the channel attention weights to obtain the feature map weighted by channel attention. The feature map weighted by channel attention is input into the spatial attention module for processing in the spatial dimension to calculate the spatial attention weights. The feature map weighted by channel attention is multiplied by the spatial attention weights to obtain the final feature map processed by CBAM. The feature map processed by CBAM is used as the input to the next layer of the convolutional neural network for subsequent feature extraction and processing.
[0039] The channel attention mechanism and the spatial attention mechanism can be expressed as:
[0040]
[0041] CBAM can perform adaptive weight assignment on the feature map in both the channel and spatial dimensions. Therefore, CBAM is embedded into the Inception module to highlight important features and suppress irrelevant features, thereby improving the performance of the model.
[0042] (2) Embed the FRN layer in the Reduction module of the Inception v4 model. Its schematic diagram is as Figure 3 shown. The overall network structure of the improved Inception v4 model is as Figure 4 shown.
[0043] The main purpose of the FRN layer is to normalize the feature map output by the convolutional layer, making the features have a more stable distribution among different filters and samples, thereby accelerating the convergence of model training and improving the generalization ability of the model. It adjusts the distribution of features by normalizing the response of each filter, reducing the problem of internal covariate shift, enabling the model to learn features more effectively.
[0044] (3) Select the Cross-Entropy Loss function on those high-quality positive examples in the training set, which helps to obtain a higher AUC value.
[0045] The loss function is used to evaluate the deviation between the predicted value and the true value in the convolutional neural network. Select the Cross-Entropy Loss function, and the formula is as follows:
[0046] The expression of the Cross-Entropy Loss function for the discrete case is:
[0047] H(p,q) = -∑p(x)log q(x)
[0048] The expression of the Cross-Entropy Loss function for the continuous case is as follows:
[0049] H(p,q) = -∫p(x)log q(x)dx
[0050] Where p(x) is the true probability, q(x) is the predicted probability of the model, and x represents different classes. For a binary classification problem, assume the true label of the positive class is 1 with probability p, and the true label of the negative class is 0 with probability 1 - p. The probability that the model predicts the positive class is The probability of the negative class is
[0051] Step 3: On the basis of the completed model construction, set specific parameters for the improved Inception v4 network. The input image size is img_size = 299×299, the initial epoch is Init_Epoch = 0, the training epoch is Epoch = 50, the maximum learning rate is Init_lr = 0.0001, the minimum learning rate is Min_lr = 0.00001, and the optimizer uses the adam optimizer. Put the improved Inception v4 network structure with the set parameters into a computer with a configured environment for training, and adjust the parameters to optimize the model, and output the image classification model after training is completed.
[0052] Step 4: Compare the improved model with the original Inception v4 model. After training with the same dataset, use the accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1_score) to compare and evaluate the models respectively.
[0053] (1) Accuracy represents the accuracy rate, which indicates the proportion of the number of samples correctly predicted by the model in the results of the algorithm's image classification, reflecting the overall prediction accuracy of the model.
[0054] The specific formula is as follows:
[0055]
[0056] (2) Precision represents the precision rate, which indicates the proportion of the true positive examples among the samples predicted as positive in the results of the algorithm's image classification.
[0057] The specific formula is as follows:
[0058]
[0059] (3) Recall represents the recall rate, which indicates the proportion of samples that are actually positive examples among the results of the algorithm's image classification and are correctly predicted as positive examples by the model. It measures the model's ability to capture positive examples.
[0060] The specific formula is as follows:
[0061]
[0062] (4) F1_score represents the F1 score, which indicates the harmonic mean of precision and recall in the results of the algorithm's image classification. It comprehensively considers precision and recall, is used to balance the relationship between the two, and more comprehensively evaluates the model performance.
[0063] The specific formula is as follows:
[0064]
[0065] Using the experimental data, the ROC curve of the model can be drawn, and the area enclosed by the curve is the AUC. This indicator is used to evaluate the performance of the model for target detection of a single category. The value range of AUC is between [values not provided]. The larger the AUC value, the better the performance of the model. When AUC = 0.5, the prediction effect of the model is equivalent to random guessing; when AUC = 1, the model can perfectly distinguish positive examples and negative examples.
[0066] Matters not covered in the invention are well-known technologies.
[0067] The above embodiments are only used to illustrate the technical concept and characteristics of the present invention, and their purpose is to enable those familiar with this technology to understand the content of the present invention and implement it accordingly, and should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. An image classification method based on an improved Inception v4 model, characterized in that: The following steps are involved: Step 1: Build a data set and divide the data, images, and labels into training set, validation set, and test set in a ratio of 8:1:
1. Step 2: Improve the Inception v4 model. The specific improvement measures are as follows: (1) The CBAM attention mechanism is embedded in the Inception structure of the Inception v4 model, combining the traditional convolution with the attention mechanism, integrating the two different paradigms into a hybrid model, reducing the computational overhead while achieving better results. (2) The FRN layer is embedded before the Reduction module of the Inception v4 model to normalize the feature map of the convolution output. This eliminates the excessive influence of certain features on the model due to the large difference in feature scales, accelerates the training speed of the model, improves the convergence efficiency, and enhances the stability and generalization ability of the model. (3) During the training process, the Cross-Entropy Loss function is used to calculate the loss by comparing it with the true label of the image; the optimizer is used to adjust the parameters to minimize the loss. Step 3: After the model is built, the training set is input into the improved network model for training, and the parameters are adjusted to optimize the model, and the trained image classification model is output. Step 4: Compare the improved model with the original Inception v4 model. Using the same data and training, use accuracy, precision, recall, and F1 score to compare and evaluate the models.
2. The image classification method based on the improved Inception v4 model according to claim 1, characterized in that: In step 2 (1), the CBAM attention mechanism processes the input feature map in the channel dimension and calculates the channel attention weight. The input feature map is multiplied by the channel attention weight to obtain the feature map weighted by the channel attention. The feature map weighted by the channel attention is input into the spatial attention module, processed in the spatial dimension, and the spatial attention weight is calculated. The feature map weighted by the channel attention is multiplied by the spatial attention weight to obtain the final feature map processed by CBAM. Finally, the feature map processed by CBAM is used as the input of the next layer of convolutional neural network to continue the subsequent feature extraction and processing.
3. The image classification method based on the improved Inception v4 model according to claim 1, characterized in that: In step 2 (2), the FRN layer normalizes the feature map output by the convolutional layer. By normalizing the response of each filter, the distribution of features is adjusted to reduce the problem of internal covariate shift, so that the model can learn features more effectively, thereby accelerating the convergence of model training and improving the generalization ability of the model.
4. The image classification method based on the improved Inception v4 model according to claim 1, characterized in that: In step 2 (3), the Cross-Entropy Loss function effectively measures the difference between the probability category predicted by the model and the true category label by continuously adjusting the parameters, guiding the model to learn the correct classification features, making the predicted distribution as close to the true distribution as possible, and improving the classification accuracy.
5. The image classification method based on the improved Inception v4 model according to claim 1, characterized in that: In step 4, the ROC curve of the model can be plotted using the experimental data. The area under the curve is the AUC (between 0.5 and 1). This indicator is used to evaluate the classification performance of the model in the binary classification problem. The closer the ROC curve is to the upper left corner, the better the performance of the model. The larger the AUC value, the better the model performance.