A Fine-Grained Image Classification Method and System Based on a Dual-Branch Network

By introducing non-maximum activation of dual-branch networks and isomorphic dual-branch subnets in fine-grained image classification, combined with the Swin Transformer architecture, the problem that convolutional neural networks cannot be extended to the Transformer architecture network is solved, and a more accurate and robust fine-grained image classification is achieved.

CN113902948BActive Publication Date: 2025-05-27ARMY ENG UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111175746.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-05-27
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

The convolutional neural network method in the prior art cannot be extended to fine-grained image classification tasks based on Transformer architecture networks, and the attention mechanism does not sufficiently extract target features.

Method used

The fine-grained image classification method based on the dual-branch network is adopted to extract the target image features through the non-maximum activation dual-branch network, including the non-maximum activation module and the isomorphic dual-branch subnet, and is improved in combination with the Swin Transformer deep neural network.

Benefits of technology

It achieves the acquisition of more sufficient target discriminant regional features, improves the accuracy and robustness of fine-grained classification results, and solves the problem that the attention mechanism method based on convolutional neural network is difficult to extend to the Transformer architecture network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902948B_ABST
    Figure CN113902948B_ABST
Patent Text Reader

Abstract

The present invention discloses a fine-grained image classification method based on a dual-branch network, including: preprocessing a target image to be classified; inputting the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; based on the obtained features of the target image to be classified, using a classifier to perform class prediction to obtain the class prediction result of the features of the target image to be classified; and using a preset fusion method to fuse the class prediction results to obtain the classification result of the target image to be classified. The present invention can solve the deficiencies that the convolutional neural network method in the prior art cannot be extended to the fine-grained image classification task based on the Transformer architecture network and the problem that the attention mechanism is insufficient in extracting target features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fine-grained image classification method and system based on a dual-branch network, belonging to the field of computer vision technology. Background Art

[0002] Fine-grained image classification belongs to the field of image classification tasks. Different from ordinary classification tasks, the fine-grained image classification task aims to classify subclasses in a large class, such as different types of cars, birds, and airplanes. Such classification targets have the characteristics of large intra-class differences and small inter-class differences, which makes the key to classification lie in the extraction of subtle features of the target. At present, the fine-grained image classification method mainly uses the attention mechanism to perform maximum activation on the target to obtain effective local discriminative features of the target, lacking the extraction of non-maximum activation features; on the other hand, most of the current fine-grained classification methods mainly extract target features based on convolutional neural networks, lacking considerations such as designing the network backbone and target function from the perspective of the Transformer architecture. The existing convolutional neural network methods are difficult to be directly extended to the fine-grained image classification method based on the Transformer architecture network. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a fine-grained image classification method and system based on a dual-branch network, which can solve the deficiencies that the convolutional neural network method in the prior art cannot be extended to the fine-grained image classification task based on the Transformer architecture network and the problem that the attention mechanism is insufficient in extracting target features. To achieve the above purpose, the present invention is implemented by the following technical solutions:

[0004] In the first aspect, the present invention provides a fine-grained image classification method based on a dual-branch network, including:

[0005] Preprocessing the target image to be classified;

[0006] Inputting the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; wherein, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network. The non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network for learning, and the features of the target image to be classified are output;

[0007] Based on the obtained features of the target image to be classified, using a classifier to perform category prediction to obtain the category prediction result of the features of the target image to be classified;

[0008] Fuse the category prediction results using a preset fusion method to obtain the classification result of the target image to be classified.

[0009] Combined with the first aspect, further, the preprocessing of the target image to be classified includes:

[0010] Scale the target image to be classified to a size of 600 pixels × 600 pixels, and crop an image area of 448 pixels × 448 pixels centered on the image center.

[0011] Combined with the first aspect, further, the non-maximum activation double-branch network includes an image preprocessing module, a backbone feature extraction module, a non-maximum activation module, and a homogeneous double-branch subnetwork.

[0012] Combined with the first aspect, further, the non-maximum activation double-branch network is trained through the following steps:

[0013] Input the images for training into the non-maximum activation double-branch network; under the guidance of a preset target loss function, use the stochastic gradient descent algorithm to train the parameters of the non-maximum activation double-branch network to obtain the optimal network parameters.

[0014] Combined with the first aspect, preferably, the images for training are obtained through the following steps: Scale the images to a size of 600 pixels × 600 pixels, randomly crop a regional image of 448 pixels × 448 pixels, and flip the images using a random horizontal flip method with a random rate of 0.5.

[0015] Combined with the first aspect, further, the preset target loss function is:

[0016] L = L CE + λL DB (1)

[0017] In formula (1), L represents the target loss function; λ represents the weight between L CE and L DB ; L CE represents the cross-entropy classification target loss function, which is expressed by the following formula:

[0018]

[0019] In formula (2), B represents the number of input image sheets; f b represents the feature of the b-th image, y b represents the true label of the b-th image, and y b ∈{1, 2, …, C}, represents the weight parameter that maps the feature f b to the true label category y b ; Wj Indicates feature f b The weight parameter mapped to the j-th category, where C is the total number of categories;

[0020] In formula (1), L DB Indicates the similarity metric objective loss function, which is expressed by the following formula:

[0021]

[0022] In formula (3), Indicates the similarity matrix between different branch features, Indicates the feature of the b-th image and The concatenated and normalized feature, Indicates The transposed feature, diag(S b ) indicates extracting the main diagonal values of S b B represents the number of input images.

[0023] Combined with the first aspect, further, the calculation process of the non-maximum activation module includes:

[0024] F m , F n = NAM(F) (4)

[0025] In formula (4), NAM(.) represents the module calculation process; F represents the feature input to the non-maximum activation module, satisfying F ∈ R B×L×C ; where B represents the number of input images, L represents the feature dimension, C represents the number of feature channels; F m Represents the maximum activation feature output by the non-maximum activation module, satisfying F m ∈ R B×L×C ,

[0026] F n Represents the non-maximum activation feature output by the non-maximum activation module, satisfying F n ∈ R B×L×C .

[0027] Combined with the first aspect, further, the calculation process of the maximum activation feature F m output by the non-maximum activation module is:

[0028] Equally divide the feature F of the input non-maximum activation module into k groups along the second dimension, requiring that the feature dimension L is an integer multiple of the number of groups k, and the i-th group of features is

[0029] Calculate the weight matrix corresponding to each group of features F i ​Calculated by the following formula:

[0030]

[0031] In formula (5), represents the weight of the l-th dimension of the b-th image, and the denominator is the weight normalization for k blocks; A i (b, l) represents the element at position (b, l) in A i ; A i represents the i-th group of eigenvectors, and is expressed by the following formula:

[0032]

[0033] In formula (6), GAP(·) represents the global average pooling operation, represents the convolution operation, and σ(·) represents the Relu activation operation;

[0034] Expand and weight the feature F in the channel dimension with the weight matrix to obtain the weighted feature i

[0035] Concatenate the k groups of features to obtain the maximum activation feature F m .

[0036] Combined with the first aspect, further, the calculation process of the non-maximum activation feature F n output by the non-maximum activation module is:

[0037] Perform the maximum suppression operation on each group of features:

[0038]

[0039] In formula (7), represents the ranking of the suppression of the weight matrix , represents the α-th largest value in, β ∈ [0, 1] represents the suppression degree of the weight of the maximum activation feature , represents the i-th group of non-maximum excitation feature weight matrices;

[0040] Expand and weight the feature F in the channel dimension with the weight matrix to obtain the weighted feature i

[0041] Concatenate the k groups of features to obtain the non-maximum activation feature F n . ​​

[0042] In combination with the first aspect, further, the step of fusing the class prediction results by using a preset fusion method includes:

[0043] The class probability prediction results are denoted as p 1 , p 2 and p 3 , corresponding to three features X 1 , X 2 and X 3 output by the non-maximum activation dual-branch network respectively; among them, X 1 represents the target image feature to be classified containing the maximum activation feature, X 2 represents the target image feature to be classified containing the non-maximum activation feature, X 3 represents the concatenated feature, satisfying X 3 =Concat(X 1 , X 2 ), and Concat(·) represents the concatenation operation in the feature channel dimension;

[0044] The prediction result is obtained by weighted sum fusion and calculated by the following formula:

[0045]

[0046] In formula (8), represents the predicted probability of the c-th class after weighted fusion, represents the probability of the c-th class output by the k-th path, C is the number of classes, and M represents the total number of paths;

[0047] The classification result of the target image to be classified is the class corresponding to the maximum value of the median, and the calculation process is

[0048] In a second aspect, the present invention provides a fine-grained image classification system based on a dual-branch network, including:

[0049] A preprocessing module: used for preprocessing the target image to be classified;

[0050] A feature extraction module: used for inputting the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the target image feature to be classified; among them, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network, the non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network for learning by the homogeneous dual-branch sub-network to output the target image feature to be classified;

[0051] Category prediction module: It is used to perform category prediction on the target image to be classified by using a classifier based on the obtained features of the target image to be classified, and obtain the category prediction result of the features of the target image to be classified;

[0052] Fusion output module: It is used to fuse the category prediction results by using a preset fusion method to obtain the classification result of the target image to be classified.

[0053] Compared with the prior art, the beneficial effects achieved by the fine-grained image classification method and system based on a dual-branch network provided by the embodiments of the present invention include:

[0054] The present invention preprocesses the target image to be classified; inputs the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; wherein, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network, the non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network, and the homogeneous dual-branch sub-network is used for learning and outputting the features of the target image to be classified; the present invention is improved on the basis of the Swin Transformer deep neural network, introducing a non-maximum activation module and a homogeneous dual-branch sub-network, which can obtain more sufficient target discriminative region features;

[0055] Based on the obtained features of the target image to be classified, the present invention uses a classifier to perform category prediction to obtain the category prediction result of the features of the target image to be classified; uses a preset fusion method to fuse the category prediction results to obtain the classification result of the target image to be classified; the present invention can solve the problem that the feature extraction of the current attention mechanism is insufficient, can solve the deficiency that it is difficult to directly extend the attention mechanism method based on the convolutional neural network to the fine-grained classification method based on the Transformer framework in the prior art, and can improve the accuracy and robustness of the fine-grained classification result. Description of the Drawings

[0056] Figure 1 is a flowchart of a fine-grained image classification method based on a dual-branch network provided by Embodiment 1 of the present invention;

[0057] Figure 2 is a structural diagram of a non-maximum activation dual-branch network in a fine-grained image classification method based on a dual-branch network provided by Embodiment 1 of the present invention;

[0058] Figure 3 is a structural diagram of a non-maximum activation module in a fine-grained image classification method based on a dual-branch network provided by Embodiment 1 of the present invention;

[0059] Figure 4It is the structural diagram of a fine-grained image classification system based on a dual-branch network provided by the second embodiment of the present invention. Detailed implementation manners

[0060] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.

[0061] Embodiment 1:

[0062] As Figure 1 shown, the embodiment of the present invention provides a fine-grained image classification method based on a dual-branch network, including:

[0063] Preprocess the target image to be classified;

[0064] Input the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; wherein, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network, and the non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network for learning by the homogeneous dual-branch sub-network to output the features of the target image to be classified;

[0065] Based on the obtained features of the target image to be classified, use a classifier to perform class prediction to obtain the class prediction result of the features of the target image to be classified;

[0066] Use a preset fusion method to fuse the class prediction results to obtain the classification result of the target image to be classified.

[0067] As Figure 1 shown, the specific steps are as follows:

[0068] Step 1: Preprocess the target image to be classified.

[0069] Scale the target image to be classified to a size of 600 pixels × 600 pixels, and crop an image area with a size of 448 pixels × 448 pixels centered on the center of the image.

[0070] Step 2: Construct and train a non-maximum activation dual-branch network.

[0071] The non-maximum activation dual-branch network includes an image preprocessing module, a backbone feature extraction module, a non-maximum activation module, and a homogeneous dual-branch sub-network.

[0072] Step 2.1: Construct a non-maximum activation dual-branch network.

[0073] After the first three stages of the existing Swin Transformer deep neural network, a non-maximum activation module is introduced, and the last stage of the Swin Transformer deep neural network is replicated to construct an isomorphic two-branch sub-network, and finally a non-maximum activation two-branch network is constructed. The constructed non-maximum activation two-branch network is as shown in Figure 2 shown below.

[0074] Specifically, the relevant content of the Swin Transformer deep neural network can be found in Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin and Baining Guo, Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, [DB][2021-04-08]https: / / arxiv.org / abs / 2103.14030.

[0075] As shown in Figure 3 shown below, the non-maximum activation module processes the output features of the first three stages to obtain the maximum activation feature and the non-maximum activation feature, and sends them to the isomorphic two-branch sub-network for processing.

[0076] The calculation process of the non-maximum activation module includes:

[0077] F m ,F n =NAM(F) (1)

[0078] In formula (1), NAM(.) represents the module calculation process; F represents the feature input to the non-maximum activation module, satisfying F ∈ R B×L×C ; where B represents the number of input image sheets, L represents the feature dimension, and C represents the number of feature channels; F m represents the maximum activation feature output by the non-maximum activation module, satisfying F m ∈ R B×L×C ,

[0079] F n represents the non-maximum activation feature output by the non-maximum activation module, satisfying F n ∈ R B×L×C .

[0080] The calculation process of the maximum activation feature F m output by the non-maximum activation module is as follows:

[0081] The feature F along the second dimension of the input non-maximum activation module is equally divided into k groups, and it is required that the feature dimension L is an integer multiple of the number of groups k, and the i-th group of features is

[0082] Calculate the weight matrix corresponding to each group of features F i Calculate through the following formula:

[0083]

[0084] In formula (2), represents the weight of the l-th dimension of the b-th image, and the denominator is the weight normalization for k blocks; A i (b, l) represents the element at position (b, l) in A i A represents the i-th group of feature vectors, which is represented by the following formula: i

[0085]

[0086] In formula (3), GAP(·) represents the global average pooling operation, represents the convolution operation, and σ(·) represents the Relu activation operation;

[0087] Expand and weight the feature F in the channel dimension with the weight matrix to obtain the weighted feature i

[0088] Concatenate the k groups of features to obtain the maximum activation feature F m .

[0089] The calculation process of the non-maximum activation feature F output by the non-maximum activation module is: n

[0090] Perform the maximum suppression operation on each group of features:

[0091]

[0092] In formula (4), represents the ranking of the suppression of the weight matrix , represents the α-th largest value in, and β ∈ [0, 1] represents the suppression degree of the maximum activation feature weight , represents the non-maximum excitation feature weight matrix of the i-th group;

[0093] Expand and weight the feature F in the channel dimension with the weight matrix i ​​​​​Perform extended weighted calculation to obtain weighted features

[0094] Concatenate k groups of features to obtain non-maximum activation feature F n .

[0095] Step 2.2: Train the non-maximum activation dual-branch network.

[0096] Input the images for training into the non-maximum activation dual-branch network; under the guidance of a preset target loss function, use the stochastic gradient descent algorithm to train the parameters of the non-maximum activation dual-branch network to obtain the optimal network parameters.

[0097] The images for training are obtained through the following steps: Scale the images to a size of 600 pixels × 600 pixels, randomly crop a region image with a size of 448 pixels × 448 pixels, and flip the images using a random horizontal flip method with a random rate of 0.5.

[0098] The preset target loss function is composed of a cross-entropy classification target loss function and a similarity metric target loss function, and is calculated by the following formula:

[0099] L = L CE + λL DB (5)

[0100] In formula (5), L represents the target loss function; λ represents the weight between L CE and L DB ; L CE represents the cross-entropy classification target loss function, and is expressed by the following formula:

[0101]

[0102] In formula (6), B represents the number of input image sheets; f b represents the feature of the b-th image, y b represents the true label of the b-th image, and y b ∈ {1, 2,..., C}, represents the weight parameter that maps the feature f b to the true label category y b , W j represents the weight parameter that maps the feature f b to the j-th category, and C is the total number of categories;

[0103] In formula (5), L DB represents the similarity metric target loss function, and is expressed by the following formula:

[0104]

[0105] In formula (7), represents the similarity matrix between different branch features, represents the feature of the b-th image, and the concatenated and normalized feature, represents the feature after transpose, diag(S b ) represents extracting the main diagonal values of S b , and B represents the number of input images.

[0106] Step 3: Input the preprocessed target image into the pre-trained non-maximum activation double-branch network to extract the target image features to be classified.

[0107] The non-maximum activation module outputs the maximum activation feature and the non-maximum activation feature, and the output features are input into the isomorphic double-branch sub-network for learning by the isomorphic double-branch sub-network to output the target image features to be classified.

[0108] Step 4: Based on the obtained target image features to be classified, use a classifier to perform class prediction to obtain the class prediction result of the target image features to be classified.

[0109] Specifically, the image I to be classified ∈ R 3×448×448 obtains three classification features X 1 , X 2 and X 3 after passing through the non-maximum activation double-branch network. For these three classification features, independent classifiers are used to perform class probability prediction to obtain the class probability prediction results p 1 , p 2 and p 3 . Among them, X 1 represents the target image feature to be classified containing the maximum activation feature, X 2 represents the target image feature to be classified containing the non-maximum activation feature, X 3 represents the concatenated feature, satisfying X 3 = Concat(X 1 , X 2 ), and Concat(·) represents the concatenation operation in the feature channel dimension.

[0110] Step 5: Use a preset fusion method to fuse the class prediction results to obtain the classification result of the target image to be classified.

[0111] Step 5.1: Based on the class probability prediction results p 1 , p 2 and p 3 , use weighted sum fusion to obtain the prediction result:

[0112]

[0113] In formula (8), represents the predicted probability of the c-th class after weighted fusion, represents the probability of the c-th class output by the k-th path, C is the number of classes, and M represents the total number of paths.

[0114] Step 5.2: The classification result of the target image to be classified is the class corresponding to the median maximum value, and the calculation process is

[0115] The present invention can obtain more sufficient target discriminative region features, can solve the problem of insufficient feature extraction of the current attention mechanism, can solve the deficiency that the attention mechanism method based on the convolutional neural network in the prior art is difficult to be directly extended to the fine-grained classification method based on the Transformer framework, and can improve the accuracy and robustness of the fine-grained classification result.

[0116] Embodiment 2:

[0117] As Figure 4 shown, the embodiment of the present invention provides a fine-grained image classification system based on a dual-branch network, including:

[0118] A preprocessing module: used for preprocessing the target image to be classified;

[0119] A feature extraction module: used for inputting the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; wherein, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network, the non-maximum activation module is used for outputting the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network for learning by the homogeneous dual-branch sub-network to output the features of the target image to be classified;

[0120] A class prediction module: used for performing class prediction on the basis of the obtained features of the target image to be classified by using a classifier to obtain the class prediction result of the features of the target image to be classified;

[0121] A fusion output module: used for fusing the class prediction results by using a preset fusion method to obtain the classification result of the target image to be classified.

[0122] Embodiment 3:

[0123] The embodiment of the present invention provides a fine-grained image classification device based on a dual-branch network, including a processor and a storage medium;

[0124] The storage medium is used for storing instructions;

[0125] The processor is configured to operate according to the instructions to execute the steps of the method described in the first embodiment.

[0126] Embodiment 4:

[0127] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first embodiment are implemented.

[0128] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in one Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0132] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A fine-grained image classification method based on a dual-branch network, characterized in that, it includes: Preprocess the target image to be classified; Input the preprocessed target image into a pre-trained non-maximum activation dual-branch network to extract the features of the target image to be classified; wherein, the non-maximum activation dual-branch network includes a non-maximum activation module and a homogeneous dual-branch sub-network, and the non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous dual-branch sub-network, and the homogeneous dual-branch sub-network is used for learning and outputting the features of the target image to be classified; Based on the obtained features of the target image to be classified, use a classifier to perform class prediction to obtain the class prediction result of the features of the target image to be classified; Use a preset fusion method to fuse the class prediction results to obtain the classification result of the target image to be classified; Among them, the step of using a preset fusion method to fuse the class prediction results includes: The class probability prediction result is denoted as p 1 , p 2 and p 3 , corresponding to the three features X 1 , X 2 and X 3 output by the non-maximum activation dual-branch network respectively; among them, X 1 represents the target image feature to be classified containing the maximum excitation feature, X 2 represents the target image feature to be classified containing the non-maximum excitation feature, X 3 represents the concatenated feature, satisfying X 3 = Concat(X 1 , X 2 ), and Concat(·) represents the concatenation operation in the feature channel dimension; Use weighted summation to fuse to obtain the prediction result, and calculate through the following formula: In Equation (8), represents the predicted probability of the c-th class after weighted fusion, represents the probability of the c-th class output by the k-th path, C is the number of classes, and M represents the total number of paths; The classification result of the target image to be classified is the category corresponding to the maximum value in, and the calculation process is 2. The fine-grained image classification method based on a dual-branch network according to claim 1, characterized in that, The preprocessing of the target image to be classified includes: Scale the target image to be classified to a size of 600 pixels × 600 pixels, and crop an image area with a size of 448 pixels × 448 pixels centered on the center of the image.

3. The fine-grained image classification method based on a dual-branch network according to claim 1, characterized in that, The non-maximum activation dual-branch network includes an image preprocessing module, a backbone feature extraction module, a non-maximum activation module and a homogeneous dual-branch sub-network.

4. The fine-grained image classification method based on a dual-branch network according to claim 1, characterized in that, The non-maximum activation dual-branch network is trained through the following steps: Input the images for training into the non-maximum activation dual-branch network; under the guidance of a preset target loss function, use the stochastic gradient descent algorithm to train the parameters of the non-maximum activation dual-branch network to obtain the optimal network parameters.

5. The fine-grained image classification method based on a dual-branch network according to claim 4, characterized in that, The preset target loss function is: Loss=L CE +λL DB (1) In Equation (1), Loss represents the target loss function; λ represents the weight between L CE and L DB ; L CE represents the cross-entropy classification target loss function, which is expressed by the following formula: In formula (2), B represents the number of input images; f b represents the feature of the b-th image, and y b represents the true label of the b-th image, and y b ∈{1, 2, …, C}, represents the weight parameter for mapping the feature f b to the true label category y b ; W j represents the weight parameter for mapping the feature f b to the j-th category, and C is the total number of categories; In formula (1), L DB represents the similarity metric objective loss function, which is expressed by the following formula: In formula (3), represents the similarity matrix between different branch features, represents the feature of the b-th image and the concatenated and normalized feature, represents the feature after transposition, diag(S b ) represents extracting the main diagonal values of S b , and B represents the number of input images.

6. The fine-grained image classification method based on a dual-branch network according to claim 1, characterized in that, The calculation process of the non-maximum activation module includes: F m ,F n = NAM(F) (4) In formula (4), NAM(.) represents the module calculation process; F represents the feature input to the non-maximum activation module, satisfying F ∈ R B ×L×C ; where B represents the number of input image sheets, L represents the feature dimension, and C represents the number of feature channels; F m represents the maximum activation feature output by the non-maximum activation module, satisfying F m ∈ R B×L×C , F n represents the non-maximum activation feature output by the non-maximum activation module, satisfying F n ∈R B×L×C .

7. The fine-grained image classification method based on a dual-branch network according to claim 6, characterized in that, The maximum activation feature F output by the non-maximum activation module m is calculated as follows: The feature F of the input non-maximum activation module is equally divided into k groups along the second dimension, and it is required that the feature dimension L is an integer multiple of the number of groups k, and the i-th group of features obtained is Calculate the feature F for each group i The corresponding weight matrix Calculate through the following formula: In formula (5), represents the weight of the l-th dimension of the b-th image, and the denominator is for weight normalization of k groups; A i (b, l) represents A i the element at position (b, l) in, A i represents the i-th group of feature vectors, which is represented by the following formula: In formula (6), GAP(·) represents the global average pooling operation, represents the convolution operation, and σ(·) represents the Relu activation operation; The weight matrix Expand and weight the feature F in the channel dimension i to obtain the weighted feature Concatenate k groups of features to obtain the maximum activation feature F m .

8. The fine-grained image classification method based on a dual-branch network according to claim 7, characterized in that, The non-maximum activation feature F output by the non-maximum activation module n is calculated as follows: Perform maximum suppression operation on each group of features: In formula (7), represents the maximum activation feature weight matrix of the i-th group of feature vectors The ranking of suppression, represents The α-th largest value in, β ∈ [0, 1] represents the maximum activation feature weight matrix of the i-th group of feature vectors The degree of suppression of, represents the non-maximum activation feature weight matrix of the i-th group of feature vectors; The non-maximum activation feature weight matrix of the i-th group of feature vectors Perform extended weighted calculation on the feature F in the channel dimension i to obtain a weighted feature Concatenate k groups of features to obtain the non-maximum activation feature F n .

9. A fine-grained image classification system based on a dual-branch network, characterized in that, it includes: A preprocessing module: used to preprocess the target image to be classified; Feature extraction module: It is used to input the preprocessed target image into a pre-trained non-maximum activation double-branch network to extract the features of the target image to be classified. Among them, the non-maximum activation double-branch network includes a non-maximum activation module and a homogeneous double-branch sub-network. The non-maximum activation module is used to output the maximum activation feature and the non-maximum activation feature, and the output features are input into the homogeneous double-branch sub-network for learning by the homogeneous double-branch sub-network to output the features of the target image to be classified. Category prediction module: It is used to perform category prediction using a classifier based on the obtained features of the target image to be classified to obtain the category prediction result of the features of the target image to be classified. Fusion output module: It is used to fuse the category prediction results using a preset fusion method to obtain the classification result of the target image to be classified. Among them, the fusion of the category prediction results using the preset fusion method includes: The class probability prediction result is denoted as p 1 , p 2 and p 3 , corresponding to the three features X 1 , X 2 and X 3 output by the non-maximum activation dual-branch network respectively; among them, X 1 represents the target image feature to be classified containing the maximum activation feature, X 2 represents the target image feature to be classified containing the non-maximum activation feature, X 3 represents the concatenated feature, satisfying X 3 = Concat(X 1 , X 2 ), and Concat(·) represents the concatenation operation in the feature channel dimension; The prediction result is obtained by weighted sum fusion and calculated by the following formula: In formula (8), represents the predicted probability of the c-th class after weighted fusion, represents the probability of the c-th class output by the k-th path, C is the number of classes, and M represents the total number of paths; The classification result of the target image to be classified is the category corresponding to the maximum value in, and the calculation process is

Citation Information

Patent Citations

  • A pedestrian abnormal behavior identification method based on 3D convolution

    CN109635790A

  • Cross-domain image classification method and system based on fine-grained domain adaptation

    CN111259941A