A fine-grained image classification method based on component mixing
By adopting component mixing strategies and mixed label calculation methods in fine-grained image classification networks, the problems of insufficient data sets and model overfitting in fine-grained image classification are solved, and higher classification accuracy and lower computing overhead are achieved.
Patent Information
- Application Number
- CN202310182745.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-03-01
AI Technical Summary
In the fine-grained image classification task, due to the small variance between classes and large intra-class variance, the data set is difficult to obtain, resulting in low model overfitting and classification accuracy.
A fine-grained image classification network based on component mixing is proposed. By positioning the component areas of an object, using a clustering algorithm to perform component areas mixing between different images, generate mixed images, and supervise the model through mixed tags to reduce the risk of overfitting.
Through component mixing strategies and mixed label calculation methods, the training data can be effectively expanded, the risk of model overfitting is reduced, classification accuracy is improved, and calculation overhead is reduced in the testing stage and inference speed is improved.
Smart Images

Figure CN116403022B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a fine-grained image classification method based on component mixing. Background Art
[0002] Image classification is one of the most important tasks in the field of computer vision. In recent years, fine-grained image classification has received more attention due to its wide range of application scenarios. For example: product classification in the smart retail scenario, where the same product of the same brand has different flavors and the visual differences between these products are relatively small, vehicle recognition in the road scenario, and biodiversity detection for organisms, etc.
[0003] Early image classification tasks focused on distinguishing coarse-grained categories, such as identifying vehicles, aircraft, birds, and insects, etc. The visual differences between coarse-grained categories are relatively large, and the recognition difficulty is relatively low. In contrast, the task of fine-grained image classification is to identify different species of birds, vehicle models, and aircraft models. Due to the small visual differences between subclasses, the between-class variance of fine-grained image classification is small. And the targets belonging to the same subclass will show relatively large differences under the influence of different lighting, poses, and occlusion factors, so the within-class variance of fine-grained image classification is large. Since it is difficult to obtain fine-grained image classification datasets, the datasets collected in actual applications are relatively few; due to the different scarcities of different subclasses, the datasets collected show a long-tailed distribution, which also exacerbates the risk of model overfitting. Summary of the Invention
[0004] In order to solve the problems existing in the prior art, the present invention proposes a fine-grained image classification network based on component mixing. First, the component regions of the object are located, and then a clustering algorithm is used on the component regions of the entire batch to perform component region mixing between different images. Through the image mixing strategy, the training data for fine-grained image classification can be effectively expanded, thereby achieving the effect of reducing the overfitting risk of the deep model.
[0005] In order to achieve the above object, the present invention proposes a fine-grained image recognition method based on component mixing, including:
[0006] (1) Obtain an image dataset and preprocess the image dataset;
[0007] (2) Build a fine-grained image classification network based on component mixing, and the fine-grained image classification network based on component mixing includes a target-level global prediction module, a component prediction module, and a component mixing module;
[0008] (3) Feed the preprocessed image dataset data into the fine-grained image recognition network based on component mixture for training to obtain a trained fine-grained image recognition network based on component mixture;
[0009] (4) Input the target image to be classified into the target-level global prediction module in the trained fine-grained image recognition network based on component mixture to obtain the classification result of the target image.
[0010] Further, the preprocessing includes:
[0011] Scale, crop, and randomly horizontally flip the images in the initial image dataset for data augmentation;
[0012] Divide the initial image dataset after data augmentation into a training set and a test set.
[0013] Further, the target-level global prediction module consists of a ResNet50 convolutional neural network, a global average pooling layer, a Softmax activation layer, and a fully connected layer;
[0014] The ResNet50 convolutional neural network serves as a feature extractor to complete feature extraction of the input image and outputs a feature map;
[0015] The feature map is input into the global average pooling layer, the Softmax activation layer, and the fully connected layer to complete classification, and the cross-entropy loss function is used for supervision.
[0016] Further, the component prediction module consists of a component detection module, a ResNet50 convolutional neural network, a global average pooling layer, a Softmax activation layer, and a fully connected layer;
[0017] The component detection module uses a feature pyramid network. Its input is the feature map output by the last convolutional layer in the target-level global prediction module, and its output is the information degree scores of a fixed number of bounding boxes and anchor box regions predicted by the feature pyramid network;
[0018] According to the bounding boxes output by the component detection module, crop and scale the input image to obtain a component image, and input the component image into the ResNet50 convolutional neural network to extract component-level convolutional features;
[0019] The feature vectors of each component image are obtained by using global average pooling on the convolutional features, and the feature vectors are input into the Softmax activation layer and the fully connected layer for classification, and the cross-entropy loss function is used for supervision.
[0020] Further, the feature pyramid network consists of three convolutional layers;
[0021] The first convolutional layer does not change the resolution of the input feature map, and the output of the first layer is the input of the second layer;
[0022] The second convolutional layer downsamples the input feature map by a factor of 2 and inputs the feature map into the third convolutional layer;
[0023] The third convolutional layer downsamples the input feature map by a factor of 2 again; each activation value of the feature maps output by these three convolutional layers respectively represents the information degree scores of the anchor box regions with different scales and different aspect ratios;
[0024] The non-maximum suppression (NMS) method is used to obtain the bounding boxes.
[0025] Furthermore, the component mixing module specifically includes:
[0026] Calculate the cosine similarity matrix of the feature vectors of the component maps extracted by the component prediction module and perform spectral clustering to obtain the clustering result of the feature vectors of the component maps;
[0027] Mix the images according to the clustering result to obtain a mixed image, and use the mixed labels to supervise the mixed image;
[0028] Input the mixed image into the ResNet50 convolutional neural network feature extractor for feature extraction, then input it into the global average pooling layer, fully connected layer and Softmax activation layer to complete classification to obtain a probability vector, and use the cross-entropy loss function for supervision.
[0029] Furthermore, the mixing of the images specifically includes:
[0030] (2.1) According to the feature vectors of the component maps Obtain the clustering matrix N:
[0031]
[0032]
[0033] N = cluster(cos(U, U T ))
[0034] where: U is the matrix composed of the feature vectors of all the component maps in a batch of images U T is the transpose matrix of U, B is the batch size, K is the number of components in each image, cos(U, U T ) is the cosine similarity matrix, and cluster() is the spectral clustering algorithm;
[0035] (2.2) Generate the corresponding hybrid mask M according to the clustering information in N and the corresponding bounding box information. The hybrid mask M is a binary mask;
[0036] (2.3) Perform pairwise mixing on a batch of images according to the mask M to generate the mixed image, and the set is denoted as
[0037] Further, the supervision of the mixed image using the mixed label specifically includes:
[0038] (1) Calculate the mixing coefficient:
[0039]
[0040]
[0041] Among them: λ a is the mixing coefficient of the input image λ b is the mixing coefficient of the input image C a,i is the area, S i is the information degree score; K is the number of components in each image;
[0042] (2) Obtain the mixed label using the obtained mixing coefficient;
[0043] Y (a,b) = (1 - λ a ) · Y a + λ b · Y b
[0044] Among them: Y a is the one-hot label of, Y b is the one-hot label of;
[0045] (3) Calculate the cross-entropy loss for the mixed image using the mixed label, and perform backpropagation on the network according to the loss value;
[0046]
[0047] Among them: is the probability vector of the mixed image.
[0048] The beneficial effects of the present invention:
[0049] 1. The present invention proposes a novel component mixing strategy, which uses weakly supervised component detection techniques and spectral clustering algorithms to find components with similar semantics of different objects, mixes components with similar semantics to generate mixed images, and enables the network to learn cross-class and cross-instance semantic component correlation information by learning the mixed images and mixed labels, thereby reducing the overfitting of the model and improving the representation ability.
[0050] 2. The present invention innovates a new method for calculating mixed labels. The advantage of this calculation method is that the label value is obtained based on the information degree score and area ratio of the mixed component graph. Compared with the recent research that uses random sampling to obtain mixed label values, the mixed label calculation method of the present invention is more interpretable. In terms of actual effects, the algorithm proposed by the present invention effectively improves the classification accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flowchart of the fine-grained image classification method based on component mixing according to an embodiment of the present invention.
[0052] Figure 2 is an overall schematic diagram of the fine-grained image classification network (training stage) based on component mixing according to an embodiment of the present invention.
[0053] Figure 3 is an overall schematic diagram of the fine-grained image classification network (testing / inference stage) based on component mixing according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be described in detail below with reference to the drawings and embodiments.
[0055] Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0056] As Figure 1 shown, an embodiment of the present invention provides a fine-grained image classification method based on image mixing, including:
[0057] S101. Obtain an image dataset and preprocess the image dataset;
[0058] In this embodiment, the bird dataset (CUB-200-2011), the car dataset (Stanford Car), and the aircraft dataset (FGVC-Aircraft) are obtained. The above three datasets are publicly available datasets for fine-grained image classification. To initially enhance the generalization ability of the fine-grained image classification model, the present invention first performs data augmentation. The specific steps are as follows: Resize to 512*512, perform central cropping with a cropping size of 448*448, and randomly horizontally flip the cropped image. According to the annotation information provided by the official, each dataset is divided into a training set and a test set. In this embodiment, a total of 5994 two-dimensional images are finally obtained as the training set.
[0059] S102. Build a fine-grained image classification network based on part mixture. The fine-grained image classification network based on part mixture includes a target-level global prediction module, a part prediction module, and a part mixture module.
[0060] For ease of understanding, as Figure 2 shown, in this embodiment, the batch size of each batch of images is fixed at 2, and 2 parts are detected for each image.
[0061] The target-level global prediction module consists of a ResNet50 convolutional neural network, a global average pooling layer, a fully connected layer, and a Softmax activation layer. Assume a batch of images is After preprocessing, it is input into the ResNet50 convolutional neural network to extract features and obtain a feature map. The feature map is input into the global average pooling layer, the fully connected layer, and the Softmax activation layer for classification to obtain a probability vector, and the cross-entropy loss function is used for supervision.
[0062] The part prediction module consists of a part detection module, a ResNet50 convolutional neural network, a global average pooling layer, a Softmax activation layer, and a fully connected layer. The part detection module uses a feature pyramid network, with the output of the last convolutional layer of ResNet50 in the target-level global prediction module as the input. For ease of description, assume here The width and height are W and H. The Feature Pyramid Network consists of three convolutional layers. The first convolutional layer does not change the resolution of the feature map. The output of the first layer is the input of the second layer. The second convolutional layer downsamples the feature map by a factor of 2, and the width and height of the output feature map are W / 2 and H / 2 respectively. The feature map is input into the third convolutional layer, and the third convolutional layer downsamples the feature map by a factor of 2 again, and the width and height of the output feature map are W / 4 and H / 4 respectively. Each activation value of the feature maps output by these three convolutional layers represents the information degree score of the anchor box regions with different scales and aspect ratios. In order to remove redundant anchor box regions, the embodiment of the present invention uses Non-maximum suppression (NMS) to retain K anchor box regions. Each anchor box information can be expressed as: {S i ,x1,y1,x2,y2}; where S i is the information degree score of the i-th component region of the input image; (x1,y1) and (x2,y2) are the coordinate information of the upper left corner and the lower right corner of the bounding box of the semantic component region respectively. According to the bounding box information of these K anchor box regions, the component map is cropped from the input image and scaled to a size of 224*224. As described above, in the embodiment of the present invention, the batch number B of each batch of input images is 2, and K is taken as 2, that is, 2 anchor box regions are retained for each input image. Therefore, there are a total of 4 component maps. Finally, the corresponding semantic component maps are cropped according to the bounding box information, denoted as where i represents the i-th input image, and j represents the j-th component map output by the component detection module. All the component maps are input into the ResNet50 convolutional neural network to extract features, and the output feature map is denoted as The feature map is input into the global average pooling layer to obtain the feature vector The feature vector corresponding to each component map is input into the fully connected layer and the Softmax activation layer for classification. The localization loss of the component detection module is shown by formulas (1) and (2):
[0063]
[0064] g(x) = max(1 - x, 0) (2)
[0065] In formula (1), P i is the probability value that the component map i is the true value class.
[0066] The component mixing module takes the feature vector in the component prediction module as the input, calculates the cosine similarity matrix and performs spectral clustering according to the similarity matrix to obtain the clustering matrix N. The specific calculation process is shown by formulas (3), (4) and (5):
[0067]
[0068]
[0069] N = cluster(cos(U, U T )) (5)
[0070] where: U is the matrix composed of all component feature vectors in a batch of images U T is the transpose matrix of U, B is the batch number, K is the number of components in each image, cos(U, U T ) is the cosine similarity matrix, and cluster() is the spectral clustering algorithm.
[0071] In this embodiment, N is a 2*2 matrix. According to the clustering information in N and the bounding box information of semantic components, we generate the corresponding hybrid mask M. The hybrid mask M is a binary mask. Based on the mask M, pairwise mixing is performed on a batch of images to generate a mixed image where a, b ∈ (1, B) and a ≠ b. It is worth mentioning that the image mixing strategy of the present invention is asymmetric.
[0072] The mixing operation can be expressed by formula (6):
[0073]
[0074] The present invention proposes a new method for calculating hybrid labels to supervise the mixed image Each semantic component map output by the component detection module corresponds to an information degree score S i , and the information degree score represents the richness of semantic information in the region. As the network is optimized, the information degree score S i predicted by the component detection module will also be dynamically optimized. Therefore, the information degree score S i can effectively describe the semantics of the component region. In addition, we also consider the spatial proportion of the component region relative to the entire component region, and use the spatial proportion of the semantic region in the entire target and the information degree score of the semantic region to determine the hybrid label value. The calculation formula of the hybrid label of the present invention is shown in formulas (7), (8) and (9):
[0075]
[0076]
[0077] Y (a,b) = (1 - λ a )·Y a + λ b ·Y b(9)
[0078] Taking formula (7) as an example, C a,i is the area of the component region i to be mixed in the a graph, and S i is the information degree score of this region, is the sum of the areas of the K component regions of the a image.
[0079] Y a and Y b are the one-hot labels of the a graph and the b graph respectively. The one-hot label refers to a one-dimensional vector in which only the true value element is 1 and the rest of the elements are 0.
[0080] The overall loss function consists of three parts: the cross-entropy loss between the original image and the semantic component graph (formula 10), the cross-entropy loss of the mixed image (formula 11), and the loss of the component detection module (formula 1). Among them, P obj and P part are the probability vectors of the input image and the component graph, and Y is the one-hot label, is the probability vector of the mixed image.
[0081] L img = -Y · log(P obj ) - Y · log(P part ) (10)
[0082]
[0083] The relevant formula of the total loss function is shown in formula (12), where α, β, and γ are all 1:
[0084] L total = αL img + βL mix + γL pair (12)
[0085] S103. Feed the data of the preprocessed image dataset into the fine-grained image recognition network based on component mixing for training to obtain a trained fine-grained image recognition network based on component mixing;
[0086] When training the classification network, set the training parameters including:
[0087] In this embodiment, the initial learning rate of the network is set to 0.001. The learning rate gradually decreases as the number of iterations increases. After 40 rounds of training, the learning rate is multiplied by 0.1, the batch size is 8, the momentum is 0.9, the weight decay is 5e-5, and the stochastic gradient descent method is used for training. The settings of the remaining network parameters can be understood conventionally and will not be elaborated here.
[0088] S104. Input the target image to be classified into the target-level global prediction module in the trained fine-grained image recognition network based on component mixing, and obtain the classification result of the target image.
[0089] As Figure 3 shown, in the test / inference stage of the present invention, the component mixing and component prediction modules are discarded. Since the feature extractor and classifier learn the cross-category and cross-instance semantic component correlation information from the mixed images in the training stage, the robustness and classification accuracy of the model are improved. Moreover, for the three modules of the present invention: the target-level global prediction module, the component prediction module, and the component mixing module, the feature extractor and the fully connected layer classifier used all share parameters. Therefore, only using the target-level global prediction module to classify the data can also ensure high accuracy in the test stage. At the same time, the computational overhead of the model is greatly reduced, and the inference speed is significantly increased.
[0090] A fine-grained image classification network based on component mixing is used in the present invention. The objective of the present invention is to solve the problems of small datasets and easy overfitting of deep convolutional models in fine-grained image classification tasks. The present invention uses a weakly supervised component detection technique to detect the semantic components of the target without any additional annotations. At the same time, the present invention proposes a novel mixing strategy based on component semantic similarity to mix the component graphs with similar semantics of different input images to generate mixed images. In addition, the present invention proposes a new calculation method for mixed labels for the generated mixed images, and supervises the model through cross-entropy loss and mixed labels, enabling the model to learn cross-category and cross-instance semantic component correlation information, thereby improving the robustness of the model and the classification accuracy. Another advantage of the present invention is that the computational overhead in the test / inference stage is very small. Only the target-level global prediction module, that is, the feature extractor plus the fully connected layer classifier, is required to complete the accurate classification of the test images. Through experiments, it is proved that the proposed fine-grained image classification network based on component mixing in the present invention has achieved state-of-the-art accuracy on three publicly available fine-grained classification datasets, as shown in Table 1.
[0091] Table 1
[0092]
[0093]
[0094] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A fine-grained image classification method based on component mixing, characterized in that, It includes the following steps: (1) Obtain an image dataset and preprocess the image dataset; (2) Build a fine-grained image classification network based on part mixing. The fine-grained image classification network based on part mixing includes a target-level global prediction module, a part prediction module, and a part mixing module; The target-level global prediction module consists of a ResNet50 convolutional neural network, a global average pooling layer, a Softmax activation layer, and a fully connected layer; The ResNet50 convolutional neural network serves as a feature extractor to complete feature extraction of the input image and outputs a feature map; The feature map is input into the global average pooling layer, the Softmax activation layer, and the fully connected layer to complete classification, and the cross-entropy loss function is used for supervision; The part prediction module consists of a part detection module, a ResNet50 convolutional neural network, a global average pooling layer, a Softmax activation layer, and a fully connected layer; The part detection module uses a feature pyramid network. Its input is the feature map output by the last convolutional layer in the target-level global prediction module, and its output is the information degree scores of a fixed number of bounding boxes and anchor box regions predicted by the feature pyramid network; According to the bounding boxes output by the part detection module, the input image is cropped and scaled to obtain part images, and the part images are input into the ResNet50 convolutional neural network to extract part-level convolutional features; The feature vectors of each part image are obtained by using global average pooling on the convolutional features, and the feature vectors are input into the Softmax activation layer and the fully connected layer for classification, and the cross-entropy loss function is used for supervision; The part mixing module specifically includes: Calculate the cosine similarity matrix of the feature vectors of the part images extracted by the part prediction module and perform spectral clustering to obtain the clustering results of the feature vectors of the part images; Mix the images according to the clustering results to obtain mixed images, and use mixed labels to supervise the mixed images; Input the mixed images into the ResNet50 convolutional neural network feature extractor for feature extraction, then input them into the global average pooling layer, the fully connected layer, and the Softmax activation layer to complete classification to obtain a probability vector, and use the cross-entropy loss function for supervision; (3) Feed the preprocessed image dataset data into the fine-grained image classification network based on part mixing for training to obtain a trained fine-grained image classification network based on part mixing; (4) Input the target image to be classified into the target-level global prediction module in the trained fine-grained image classification network based on part mixing to obtain the classification result of the target image.
2. The fine-grained image classification method based on component mixing according to claim 1, wherein The preprocessing includes: Scale, crop, and randomly horizontally flip the images in the initial image dataset for data augmentation; divide the augmented initial image dataset into a training set and a test set.
3. The fine-grained image classification method based on component mixing according to claim 1, characterized in that: The feature pyramid network consists of three convolutional layers; The first convolutional layer does not change the resolution of the input feature map, and the output of the first layer is the input of the second layer; The second convolutional layer downsamples the input feature map by a factor of 2 and inputs the feature map into the third convolutional layer; The third convolutional layer downsamples the input feature map by a factor of 2 again; each activation value of the feature maps output by these three convolutional layers respectively represents the information degree scores of the anchor box regions with different scales and different aspect ratios; The non-maximum suppression (NMS) method is used to obtain the bounding boxes.
4. The fine-grained image classification method based on component mixing according to claim 1, characterized in that, The mixing of the images specifically includes: (2.1) According to the eigenvector of the component diagram Obtain the clustering matrix N: N = cluster(cos(U,U T )) Where: U is the feature vector of all the component diagrams in a batch of images The matrix composed of, U T Is the transpose matrix of U, B is the batch quantity, K is the number of components in each diagram, cos(U, U T ) is the cosine similarity matrix, and cluster() is the spectral clustering algorithm; (2.2) Generate a corresponding mixing mask M according to the clustering information in N and the corresponding bounding box information, and the mixing mask M is a binary mask; (2.3) Pairwise mixing is performed on a batch of images according to the mask M to generate the mixed images, and the set thereof is denoted as 5. The fine-grained image classification method based on component mixing according to claim 1, characterized in that, The supervision of the mixed image using the mixed labels specifically includes: (1) Calculate the mixing coefficient: Where: λ a is the input image 's mixing coefficient, λ b is the input image 's mixing coefficient; C a,i is 's area, S i is 's information degree score; K is the number of components in each image; (2) Obtain the mixed labels using the obtained mixing coefficient; Y (a,b) = (1 - λ a ) · Y a + λ b · Y b Wherein: Y a is 's one-hot label, and Y b is 's one-hot label; (3) Calculate the cross-entropy loss for the mixed image using the mixed labels, and perform backpropagation on the network according to the loss value. Wherein: is the probability vector of the mixed image.
Citation Information
Patent Citations
Weak supervision fine-grained image classification method of multi-branch neural network model
CN111178432A
Fine-grained image classification method fusing multi-granularity features
CN113688894A