Bird fine-grained image recognition method and device based on multi-scale feature extraction and region alignment
Through the multi-scale feature extraction and regional alignment methods, the problem of difficulty in effectively fusion of global and local features in bird image recognition is solved, and efficient recognition of bird fine-grained images is achieved, which improves the robustness and accuracy of the recognition.
Patent Information
- Application Number
- CN202510408257.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively extract global features that reflect the overall spatial structure relationship and capture local features that capture subtle differences, resulting in unstable classification performance in fine-grained image recognition tasks, especially in bird image recognition, global features are difficult to capture fine-grained differences, while local features are susceptible to interference and occlusion.
The multi-scale feature extraction and regional alignment method is adopted to extract global features through the backbone network and enhance it using spatial attention mechanism. Combined with dynamic anchor boxes, local key areas are generated, local features are aligned and regularized, and finally the integration of global features and local features is achieved through self-supervised alignment and regularization strategies.
It improves the robustness and discriminant ability of bird fine-grained image recognition, and can retain overall spatial semantic stability while strengthening the sensitivity to fine-grained difference, improving the accuracy of image recognition.
Smart Images

Figure CN120260079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and designs a fine-grained bird image recognition method and device for multi-scale feature extraction and region alignment. Background Art
[0002] There are many objects in the real world that are extremely similar visually, belong to the same large category, but have subtle differences and belong to different sub-categories. For example, the appearance forms of Chinese roses, roses, and multiflora roses are extremely similar, and it is easy to confuse them as the same kind of flower, but these three kinds of flowers are actually different varieties under the genus Rosa. The fine-grained image recognition task is to let the machine help humans distinguish objects that are highly similar in overall appearance but belong to different categories. Compared with traditional image recognition tasks, fine-grained image recognition pays more attention to mining the subtle differences that can distinguish different sub-categories, and this kind of subtle difference is also called the discriminative region. How to extract both the global features reflecting the overall spatial structure relationship and the local features containing subtle differences is the key problem in fine-grained image recognition.
[0003] Deep convolutional neural networks can capture the spatial layout of objects and the relative position relationships between components through their hierarchical receptive fields. Such global spatial features have strong robustness and can, to a certain extent, offset the instability caused by local detail loss and noise interference, providing stable class clues for the model. However, it is difficult for global information to capture the tiny differences between fine-grained classes, which is not enough to distinguish highly similar classes for the fine-grained image recognition task. To solve this problem, local features need to be introduced to focus on the key region information in the image. However, directly extracting local features is often affected by background interference, local noise, local occlusion, and deformation, which easily introduces redundant or incorrect information and reduces the overall classification performance. And there may be problems with difficult local feature alignment due to changes in spatial relationships (object pose changes and camera perspective changes) in the same part between different samples. In addition, separate local features lack the constraint of global spatial relationships and are prone to the "fragmentation" phenomenon, that is, although local features have high discriminability, due to the lack of overall spatial context information, it may be difficult to support the final classification. Therefore, effectively fusing global features and local features can, while retaining the stability of the overall spatial semantics, also enhance the model's sensitivity to fine-grained differences. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a fine-grained image recognition method and device combining multi-scale feature extraction and key area alignment, extracting global features and local features from the global and local branches respectively. The global branch ensures the capture of the overall semantics of the image, and the local branch finely captures the local area with discriminative features through key area alignment and regularization. The complementary fusion of the two makes the final feature representation both robust and has high fine-grained discriminative ability.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] S1: Collect bird images and build training and testing datasets;
[0007] S2: Construct a fine-grained bird image recognition network with multi-scale feature extraction and region alignment;
[0008] S3: After preprocessing, the bird image is input into the constructed model for training;
[0009] S4: Use the bird fine-grained image recognition network with multi-scale feature extraction and region alignment to apply the trained model to actual bird image recognition;
[0010] Further, the step S2 specifically includes the following steps:
[0011] S21: Use the pre-trained backbone network to extract multi-scale features from a batch of input images I, and use the spatial attention mechanism to further enhance the features extracted by the backbone network;
[0012] S22: Embed the branch activation structure and self-attention feature representation mechanism in the backbone network to perform global feature aggregation, enhancing the multi-scale perception and feature fusion capabilities of the backbone network;
[0013] S23: Use the region proposal network based on dynamic anchor boxes to screen local key areas, and use the backbone network to extract local area features for the second time;
[0014] S24: Align the local key regions to generate stable key region feature expressions;
[0015] S25: Calculate the loss of each branch and the total loss;
[0016] Further, the step S21 specifically includes the following steps:
[0017] S211: Return the intermediate features X={x1,...,x1,...,x2,...,x3,...,x4,...,x5,...,x6,...,x7,...,x8,...,x9,...,x10,...,x11,...,x12,...,x13,...,x14,...,x15,...,x16,...,x17,...,x18,...,x19,...,x219, m ,...,x n}, x mThe feature map representing the output of the last convolutional layer of the m-th main layer, where n represents the number of main layers in the backbone network;
[0018] S212: Obtain the feature maps Feature = (fs1, fs2) output by the last two convolutional layers of the last module of the last main layer of the backbone network;
[0019] S213: Use the feature fs1 as an attention map to perform spatial attention weighting on the feature fs2 through inner product calculation. First, flatten the spatial dimensions of fs1 and fs2 into one dimension, and use Einstein summation convention to calculate the tensor multiplication of fs1 and fs2 in the spatial dimension to obtain φ sp ;
[0020] S214: Divide φ sp by the number of channels of the attention tensor and perform a square root operation, then perform L2 normalization on φ sp along the channel dimension. Adjust the final result back to the spatial dimension of the original feature fs2 to obtain the feature sp_feature;
[0021] Furthermore, the specific steps in step S22 include the following steps:
[0022] S221: Perform convolution operations on the feature x with the shape of [B, C, H, W] output by the last layer of the backbone network through multiple parallel branches (each branch is a 1×1 convolutional layer), reducing the number of channels of the feature to 1, and generating multiple activation maps Activations = {A1,..., A n ,..., A m ,..., A n};
[0023] A m = branch m (x n )
[0024] A m represents the activation map generated by the m-th branch processing the feature x n , and n represents the total number of branches. These activation maps reflect the attention degree of the model to different regions, similar to the attention mechanism, representing the importance of each spatial position.
[0025] S222: Concatenate the activation maps together to obtain an activation map Activation with the dimension of [B, n, H, W]. To achieve attention to important regions, unfold the activation map Activation in the spatial dimension, and its shape becomes [B, n, H×W].
[0026] S223: Normalize each spatial position of the activation map Activation using the softmax function to obtain a weighted activation map. Then, perform a rearrangement operation on the activation map to change its shape to [B, H×W, n].
[0027] S224: Rearrange the original feature map x n to change its shape to [B, C, H×W]. Weight x n using the activation map Activation, that is, multiply the feature value at each position of x n by the corresponding activation weight at that position. After weighting, the shape of the resulting feature N is [B, n, C].
[0028] N = x n × softmax(Activation)
[0029] S225: Change the shape of the feature N to [B, C, n] through a rearrangement operation.
[0030] S226: Perform dimensionality reduction and non - linear mapping on the node feature N through a multi - layer perceptron to adjust it to a dimension suitable for the input of the self - attention feature representation mechanism.
[0031] S227: Treat the processed feature N as a sequence and input it into the self - attention feature representation mechanism for global interaction and aggregation;
[0032] N attention = AttentionEncoder(MLP(N))
[0033] Furthermore, the specific steps in step S23 include the following steps:
[0034] S231: Generate three feature maps with different resolutions from sp_feature through three - layer convolution to capture multi - scale target information;
[0035]
[0036] S232: Predict the original anchor box parameters P m at each position from the feature map sp_feature′ m generated by the m - th layer convolution through 1×1 convolution, and then convert it into size s m and aspect ratio r m ;
[0037] P m = Conv 1×1 (sp_feature′ m )
[0038] s b,i,j,k,m = Softplus(p b,i,j,k,m ) + 10 -3
[0039] r b,i,j,k,m = Sigmoid(p b,i,j,k+K,m ) × 2 + 0.5
[0040] where P m is the original parameter tensor predicted by the m-th layer, with a shape of [B, 2K, H m , W m ; K is the number of anchor boxes at each position; s b,i,j,k,m is the size of the k-th anchor box at the position (i, j) of the feature generated by the m-th layer convolution of the b-th sample; r b,i,j,k,m is the aspect ratio of the k-th anchor box at the position (i, j) generated by the m-th layer convolution of the b-th sample; p b,i,j,k,m is the original value corresponding to the size in P m ; p b,i,j,k+K,m is the original value corresponding to the aspect ratio.
[0041] S233: Generate the base anchor boxes by combining the predicted size s m and the aspect ratio r m with the center coordinates of the feature map grid;
[0042] S234: Predict the confidence and offset of each anchor box from the feature map sp_feature' m ;
[0043]
[0044] where S m is the confidence of the m-th layer, with the shape adjusted to [B, H m × W m × K]; D m is the offset of the m-th layer, with the shape adjusted to [B, H m × W m × K, 4].
[0045] S235: Apply the offset predicted for each layer to the base anchor boxes to generate the final bounding boxes B;
[0046] S236: Select the top n regions through non-maximum suppression, and the confidence scores of these regions are stored in the list score;
[0047] S237: Crop the region I part corresponding to the bounding box B from the input image I, and unify the size to 224 × 224 through bilinear interpolation;
[0048] Ipart = Interpolate(I[:, :, y′ 1,m :y′ 2,m , x′ 1,m :x′ 2,m , (224, 224))
[0049] where (x′ 1,m , y′ 1,m ) and (x′ 2,m , y′ 2,m ) represent the upper - left and lower - right coordinates of the boundary predicted by the m - th layer of convolution.
[0050] S238: Re - feed the cropped image I part into the backbone network to extract the features of the local key regions. When extracting features, the same hierarchical feature extraction idea as the global branch is adopted, and the local region features FP = {fp1,..., fp m ,..., fp n} are extracted layer - by - layer from the backbone network, where the value range of m is from 1 to topn;
[0051] S24: To solve the inconsistency of local features caused by pose and perspective changes, a key region alignment module is adopted to re - order the local features {fp1,..., fp m ,..., fp n} to form a stable key region feature representation.
[0052] Furthermore, the step S24 specifically includes the following steps:
[0053] S241: Generate a random center when there is no initial sample in the initialization stage;
[0054]
[0055] where C represents the cluster center matrix, with dimensions K×D, K representing the number of cluster centers, and D representing the feature dimension.
[0056] S242: Create a backup feature center dictionary to back up each cluster center to ensure that there is historical information available for backtracking during the update process.
[0057] C backup,k = C k , k = 0, 1,..., K - 1
[0058] where C backup,k represents the k - th backup cluster center, and C k represents the k - th initial cluster center.
[0059] S243: Calculate the feature fp of each critical region in the initial stage m to the Euclidean distance to the cluster center k;
[0060] D b,m,k =‖P b,m -C k ‖2, D ∈ R B×topn×K
[0061] where D b,m,k represents the Euclidean distance from the m-th critical region in batch b to the k-th cluster center, and P m,n represents the feature vector of the m-th critical region in batch b.
[0062] S244: Subsequently, take the index of the nearest center as the initial cluster label;
[0063]
[0064] where O b,m represents the index that finds the minimum Euclidean distance among all cluster centers for the m-th critical region in batch b.
[0065] S245: Calculate the similarity between pairwise cluster centers using matrix multiplication, and form the relationship matrix A between the centers after normalization center . Calculate the pairwise similarity for all critical regions within each image, and also perform normalization to obtain the relationship matrix A;
[0066] S246: Construct a cost matrix M by multiplying the cluster center matrix and the critical region matrix and taking the negative value, and then use the linear programming method to find the optimal matching O b of batch b, obtain the optimal permutation index of the critical regions of each image, so as to ensure that similar critical regions have a consistent order in different images;
[0067] S247: When training the images of the next batch, each cluster center will be updated according to the critical regions assigned to this center in the current batch. First, calculate the Euclidean distance between each critical region in the current batch and its corresponding cluster center. Use the exponential function to convert the distance into a weight, and the smaller the distance, the higher the weight. Normalize the weights within each sample to obtain w b,m , ensuring that the sum of the weights of all critical performance regions is 1.
[0068]
[0069] where w b,m represents the normalized weight of the m-th critical region in batch b; represents the cluster center to which the m-th critical region is assigned, O b,mis the corresponding index;
[0070] Select all the key regions assigned to the cluster center k, and use the corresponding weight w b,m to multiply with the feature vector P of this key region b,m and then obtain the weighted feature sum of the cluster center k through summation. Then divide the weighted feature sum by the total weight of all the key regions assigned to the cluster center k to obtain the weighted average;
[0071]
[0072] where C′ k represents the weighted average center of the k-th cluster.
[0073] To maintain the smoothness of the update, a smoothing factor α is introduced to control the degree of fusion between the new and old centers (C′ k is the new center calculated in this round, and C k is the old center obtained in the previous round). This mechanism of smooth update not only retains the information of the previous cluster center but also can adjust the center position in a timely manner according to the data of the new batch;
[0074] C k = αC k + (1 - α)C′ k
[0075] where C k represents the updated cluster center.
[0076] S248: Obtain the optimal matching order O of the current batch through the key region alignment module b After that, reorder the features fp1, fp2, and fp3 to obtain fm1, fm2, and fm3;
[0077] Furthermore, the specific steps in the step S25 include the following steps:
[0078] S251: The global feature branch processes the obtained global features f1, f2, and sp_feature through 3 different convolutional modules respectively. These 3 convolutional modules adopt a similar module combination, that is, first perform channel mapping by a 1×1 convolutional module, then extract local features by a 3×3 convolutional module, and finally compress the feature map into a 1×1 vector representation through an adaptive pooling layer. Then send them into different classifiers for classification to obtain the results y1, y2, and y3. After splicing f1, f2, and sp_feature, perform joint classification to obtain the classification result y4.
[0079] S252: The node features N obtained by the auxiliary branch attention also use a classifier for prediction to obtain the prediction result y5.
[0080] S253: Calculate the losses for features y1, y2, y3, y4, and y5 respectively using smooth cross-entropy loss to obtain (loss1, loss2, loss3, loss4, loss5), and add these losses to get L global :
[0081]
[0082] where p i represents the smoothing parameter to balance the contributions of predictions from each layer; z represents the true class label.
[0083] S254: The processing of fp1, fp2, and fp3 is similar to that of the global features. They are processed through a series of convolutional layers and adaptive pooling layers, and then sent to different classifiers for classification to obtain classification results yp1, yp2, and yp3. Use a concatenated classifier to concatenate fp1, fp2, and fp3 and perform joint classification to obtain classification result yp4.
[0084] S255: Use the cross-entropy loss function to calculate the losses of these classification results to obtain (losspart1, losspart2, losspart3, losspart4).
[0085] S256: Based on the predicted value yp4 of the key region and the true label, calculate the classification loss of each region through negative log-likelihood loss. The ranking loss function uses this loss as the ranking basis to force the model to align the confidence score score of the key region with the classification performance, generating a ranking loss loss_key. When the model assigns a higher confidence score to a region with poor classification performance, this loss punishes it, thus prompting the model to learn that the ranking score of important regions is higher than that of unimportant regions.
[0086] S257: Add losspart1, losspart2, losspart3, losspart4, and loss_key to get L local ;
[0087] S258: Regularize the sorted key region features fm1, fm2, and fm3 to ensure that the features after passing through the key region alignment module still retain the global semantic consistency of the local features. Calculate the distribution difference between the regularized features fm1, fm2, and fm3 and the global features f1, f2, and sp_feature through the regularization loss, and accumulate the obtained results to get L reg ;
[0088] S259: Calculate the loss L of the entire network;
[0089] L = L local + L global + L reg
[0090] The present invention also provides a fine-grained bird image recognition device for multi-scale feature extraction and region alignment, and the device includes:
[0091] An image acquisition module for acquiring bird images; the bird images are the bird images to be recognized for fine-grained bird images;
[0092] An image preprocessing module for preprocessing the bird images to generate standard bird images;
[0093] An image recognition module for inputting the standard bird images into a pre-trained network model to predict the results of fine-grained bird image recognition.
[0094] The beneficial effects of the present invention are as follows: The present invention proposes a method and device for fine-grained bird image recognition with multi-scale feature extraction and region alignment. The global feature extraction branch of the present invention performs multi-level feature extraction on the input image through a backbone network, and uses an auxiliary branch to enhance the feature expression ability of the backbone network for local region information and global context information. The local feature extraction branch of the present invention uses a region proposal network based on dynamic anchor boxes to generate local candidate regions, and then sends the cropped local region images into the backbone network again to extract key region features. Key region alignment is used to generate stable key region features. Finally, during the training process, the fusion of global features and local features is achieved through self-supervised alignment and regularization strategies.
[0095] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Description of the Drawings
[0096] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings, where:
[0097] Figure 1 It is a schematic flowchart of a method and device for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to the present invention.
[0098] Figure 2 It is a structural diagram of a method for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to the present invention.
[0099] Figure 3 This is the structural diagram of the auxiliary branch in a fine-grained bird image recognition method for multi-scale feature extraction and region alignment according to the present invention. Specific embodiments
[0100] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0101] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.
[0102] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or position relationship, they are based on the orientation or position relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the position relationship in the attached drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0103] The present invention provides a fine-grained image recognition method and device combining multi-scale feature extraction and key region alignment, including the following steps:
[0104] S1: Collect bird images and construct training and test data sets;
[0105] S2: Construct a fine-grained bird image recognition network for multi-scale feature extraction and region alignment;
[0106] S3: After preprocessing the bird images, input them into the constructed model for training;
[0107] S4: Use the fine-grained bird image recognition network for multi-scale feature extraction and region alignment to apply the trained model to actual bird image recognition;
[0108] Furthermore, the specific steps in step S1 include the following steps:
[0109] S11: Bird images can be obtained by means of public datasets, web crawling, or self-shot. To ensure the integrity and balance of the data, the collected data needs to be screened to remove blurred, duplicate, or irrelevant pictures, and ensure that the number of samples in each category is roughly balanced;
[0110] S12: Organize the data folder structure in a standard format, store pictures by category, and create relevant annotation files at the same time, such as the mapping relationship between pictures and categories, training / test data division files, category name lists, etc. The number of training set and test set images is divided according to a ratio of 5 to 1.
[0111] S13: Adjust the picture size to 448×448, remove noise, perform color normalization, and at the same time perform data augmentation, including random flipping, rotation, and color jitter, to increase data diversity and improve the generalization ability of the model.
[0112] Furthermore, the specific steps in step S2 include the following steps:
[0113] S21: Use a pre-trained backbone network to perform multi-scale feature extraction on a batch of input images I, and use a spatial attention mechanism to further enhance the features extracted by the backbone network;
[0114] S22: Embed a branch activation structure and a self-attention feature representation mechanism in the backbone network for global feature aggregation, enhancing the multi-scale perception ability and feature fusion ability of the backbone network;
[0115] S23: Use a region proposal network based on dynamic anchor boxes to screen local key regions, and use the backbone network to extract local region features {fp1, fp2, fp3} again;
[0116] S24: Input the local feature fp3 into the key region alignment module to obtain the optimal arrangement order O b , generating a stable key region feature representation;
[0117] S25: Calculate the losses of each branch and the total loss;
[0118] Furthermore, the specific steps in step S21 include the following steps:
[0119] S211: Use ResNet-50 as the backbone network to extract image features. Extract the features f1, f2, and f3 output by the last three layers of ResNet-50. The shapes of the features are [16, 512, 56, 56], [16, 1024, 28, 28], and [16, 2028, 14, 14], respectively.
[0120] S212: Obtain the feature maps Feature=(fs1, fs2) output by the last two convolutional layers in the last residual module of the last main layer of ResNet-50. The feature dimensions are [16, 512, 14, 14] and [16, 2048, 14, 14], respectively.
[0121] S213: Use the feature fs1 as the attention map to perform spatial attention weighting on the feature fs2 through inner product calculation. Use a convolutional module to adjust fs1 and fs2 to the same dimension, and the shape becomes [16, 512, 196]. Use Einstein summation convention to calculate the tensor multiplication of fs1 and fs2 in the spatial dimension to obtain φ sp , with a shape of [16, 512, 196].
[0122] S214: Divide φ sp by the number of channels of the attention tensor and perform a square root operation. Then, perform L2 normalization on φ sp along the channel dimension and reconstruct the spatial dimension to [16, 196, 14, 14]. Adjust the finally obtained result back to the spatial dimension of the original feature fs2 to obtain the feature sp_feature, and reshape the dimension to [16, 2048, 14, 14].
[0123] The step S22 specifically includes the following steps.
[0124] S221: Perform convolution operations on f3 through 10 parallel 1×1 convolutional layers to reduce the number of channels of the feature to 1, and obtain 10 activation maps Activations={A1,...,A m ,...,A 10} with a shape of [16, 1, 14, 14].
[0125] S222: Concatenate the activation maps Activations to obtain Activation with a shape of [16, 10, 14, 14]. To achieve attention to important regions, unfold the activation map Activation in the spatial dimension, with a shape of [16, 10, 196].
[0126] S223: Normalize each spatial position of the activation map Activation using the Softmax function to obtain a weighted activation map. Then, perform a rearrangement operation on the activation map to make its shape [16, 196, 10].
[0127] S224: Rearrange f3 to make its shape [16, 2048, 196]. Use the activation map Activation to perform a weighted calculation on f3 to obtain the feature N, with a shape of [16, 2048, 10].
[0128] N = x n × softmax(Activation)
[0129] S225: Rearrange the feature N to make its shape [16, 10, 2048].
[0130] S226: Define a Transformer encoder layer with an input feature dimension of 200 and 4 attention heads. Subsequently, stack this encoder layer twice to obtain a Transformer encoding module. Reshape the shape of N to [10, 16, 200] to adapt to the input format of the Transformer encoding module, and perform global feature modeling through the Transformer. Finally, restore the output N attention to its original shape [16, 10, 200].
[0131] N attention = AttentionEncoder(MLP(N))
[0132] Furthermore, the specific steps in step S23 include the following steps:
[0133] S231: Generate feature maps with different resolutions through three layers of 3×3 convolutions, with shapes [16, 2048, 14, 14], [16, 2048, 7, 7], and [16, 2048, 4, 4] respectively.
[0134]
[0135] S232: Predict the original anchor box parameters at each position from the 3 feature maps through a 1×1 convolutional layer, and then convert them into the size s and aspect ratio r. Generate 9 anchor boxes at each position, and each anchor box has 2 parameters. Therefore, the shapes of the original anchor box parameters for the three feature maps with different scales are [16, 18, 14, 14], [16, 18, 7, 7], and [16, 18, 4, 4].
[0136] P m = Conv 1×1(sp_feature′ m )
[0137] s b,i,j,k,m = Softplus(p b,i,j,k,m ) + 10 -3
[0138] r b,i,j,k,m = Sigmoid(p b,i,j,k+K,m ) × 2 + 0.5
[0139] where P m is the original parameter tensor predicted by the m-th layer, with shape [B, 2K, H m , W m ; K is the number of anchor boxes at each position; s b,i,j,k,m is the k-th anchor box size of the feature generated by the m-th layer convolution of the b-th sample at position (i, j); r b,i,j,k,m is the aspect ratio of the k-th anchor box generated by the m-th layer convolution of the b-th sample at position (i, j); p b,i,j,k,m is the original value corresponding to the size in P m ; p b,i,j,k+K,m corresponds to the original value of the aspect ratio.
[0140] S233: Generate the base anchor boxes by combining the predicted size s m and aspect ratio r m with the center coordinates of the feature map grid.
[0141] S234: Predict the confidence and offset of each anchor box from the feature map sp_feature′ m . The confidence dimensions predicted by different scale feature maps are [16, 9, 14, 14], [16, 9, 7, 7], and [16, 9, 4, 4]. After adjusting the shape, they become [16, 1764], [16, 441], and [16, 144]. Then, the confidences of all scales are concatenated together, and the final shape is [16, 2349]. The offset dimensions predicted by different scale feature maps are [16, 36, 14, 14], [16, 36, 7, 7], and [16, 36, 4, 4]. After adjusting the shape, the dimensions become [16, 1764, 4], [16, 441, 4], and [16, 144, 4]. Then, the offsets of all scales are concatenated together, and the final shape is [16, 2349, 4].
[0142]
[0143] where S m is the confidence of the m-th layer, and after adjusting the shape, it is [B, H m × W m × K]; D mis the offset of the m-th layer, and the adjusted shape is [B, H m ×W m ×K, 4].
[0144] S235: Apply the offset predicted for each layer to the base anchor boxes to generate the final bounding boxes B, with dimensions [16, 2349, 4].
[0145] S236: Select the top 4 regions through non-maximum suppression. The dimensions of these 4 regions are [16, 4, 4]. The confidence scores are stored in the list score.
[0146] S237: Crop the regions corresponding to the bounding boxes B from the input image I to obtain a set of local region images. Resize the images to 224×224 through bilinear interpolation. Finally, the dimensions of this set of images are [16, 4, 3, 224, 224].
[0147] S238: Feed the cropped image I part back into the backbone network to extract the features of the local key regions. When extracting features, use the same hierarchical feature extraction idea as the global branch, and layer by layer extract the local region features FP = {fp1, fp2, fp3} from the backbone network, with dimensions [64, 512, 28, 28], [64, 512, 28, 28], and [64, 2048, 7, 7] respectively.
[0148] Furthermore, the specific steps in step S24 include the following steps:
[0149] S241: In the case of no initial samples in the initialization stage, generate random centers, with dimensions [4, 1024];
[0150]
[0151] where C represents the cluster center matrix, with dimensions K×D; K represents the number of cluster centers, taking the value 4; D represents the feature dimension.
[0152] S242: Create a backup feature center, with dimensions [4, 1024].
[0153] C backup,k = C k , k = 0, 1,..., K - 1
[0154] where C backup,k represents the k-th backup cluster center, and C k represents the k-th initial cluster center.
[0155] S243: Resize the dimension of fp3 to [16, 4, 1024], calculate the Euclidean distance D from each critical region feature fp3 to the clustering center k, with the dimension of [16, 4, 4];
[0156] D b,m,k =‖P b,m -C k ‖2, D ∈ R B×topn×K
[0157] where D b,m,k represents the Euclidean distance from the m-th critical region in batch b to the k-th clustering center, and P m,n represents the feature vector of the m-th critical region in batch b.
[0158] S244: Then take the index of the nearest center as the initial clustering label.
[0159]
[0160] where O b,m represents the index with the minimum Euclidean distance among all clustering centers for the m-th critical region in batch b, with the dimension of [16, 4].
[0161] S245: Calculate the similarity between pairwise clustering centers using matrix multiplication, and form the relationship matrix A center , with the dimension of [4, 4]. Calculate the pairwise similarity for all critical regions within each image, and also normalize it to obtain the relationship matrix A, with the dimension of [16, 4, 4].
[0162] S246: Construct a cost matrix M with the dimension of [16, 4, 4] by multiplying the clustering center matrix and the critical region matrix and taking the negative value. Then use the linear programming method to find the optimal matching O b of batch b, and obtain the optimal permutation index of the critical regions for each image, so as to ensure that similar critical regions have a consistent order in different images. The dimension of O b is [16, 4].
[0163] S247: When training the images of the next batch, each clustering center will be updated according to the critical regions assigned to this center in the current batch. First, calculate the Euclidean distance between each critical region in the current batch and its corresponding clustering center. Use the exponential function to convert the distance into a weight, and the smaller the distance, the higher the weight. Normalize the weights within each sample to obtain w b,m , ensuring that the sum of the weights of all critical performance regions is 1.
[0164]
[0165] where w b,m represents the normalized weight of the m-th critical region in batch b, with a dimension of [16, 4]; represents the cluster center to which the m-th critical region is assigned, and O b,m is the corresponding index;
[0166] Select all the critical regions assigned to the cluster center k, and use the corresponding weight w b,m to multiply with the feature vector P b,m of this critical region, and then obtain the weighted feature sum of the cluster center k through summation. Then divide the weighted feature sum by the total weight of all the critical regions assigned to the cluster center k to obtain the weighted average;
[0167]
[0168] where C′ k represents the weighted average center of the k-th cluster.
[0169] To maintain the smoothness of the update, a smoothing factor α is introduced to control the degree of fusion of the old and new centers. This mechanism of smooth update not only retains the information of the previous cluster center but also can adjust the center position in a timely manner according to the data of the new batch;
[0170] C k = αC k + (1 - α)C′ k
[0171] where C k represents the updated cluster center.
[0172] S248: Obtain the optimal matching order O of the current batch through the critical region alignment module b After that, reorder the features fp1, fp2, and fp3 to obtain fm1, fm2, and fm3, with a dimension of [16, 4, 1024].
[0173] Furthermore, the specific steps in step S25 include the following steps:
[0174] S251: The global feature branch processes the obtained global features f1, f2, and sp_feature through 3 different convolutional modules respectively. These 3 convolutional modules adopt a similar module combination, that is, first perform channel mapping by a 1×1 convolutional module, then extract local features by a 3×3 convolutional module, and finally compress the feature map into a 1×1 vector representation through an adaptive pooling layer. The feature dimensions become [16, 256], [16, 512], and [16, 1024] respectively. Then send them into different classifiers for classification to obtain the results y1, y2, and y3, with a dimension of [16, 200].
[0175] S252: Concatenate f1, f2, and sp_feature and then perform joint classification to obtain the classification result y4, with a dimension of [16, 200]. The node feature N obtained by the auxiliary branch attention Similarly, use the classifier to make predictions to obtain the prediction result y5, with a dimension of [16, 200].
[0176] S253: Use smooth cross-entropy loss to calculate the losses of features y1, y2, y3, y4, and y5 respectively to obtain (loss1, loss2, loss3, loss4, loss5), and add these losses to obtain L global . Different smoothing factors p = {0.7, 0.8, 0.9, 1, 1} are used during loss calculation to balance the contributions of predictions at each layer. In the formula, z represents the true class label.
[0177] The processing of fp1, fp2, and fp3 is similar to that of the global feature. After being processed by a series of convolutional layers and adaptive pooling layers, the dimension becomes [64, 1024], and then they are sent to different classifiers for classification to obtain the classification results yp1, yp2, and yp3, with the dimension becoming [64, 200]. Use a concatenated classifier to concatenate fp1, fp2, and fp3 and then perform joint classification to obtain the classification result yp4, with a dimension of [64, 200].
[0178] S255: Use the cross-entropy loss function to calculate the losses of these classification results to obtain (losspart1, losspart2, losspart3, losspart4).
[0179] S256: Based on the predicted value yp4 of the key region and the true label, calculate the classification loss of each region through the negative log-likelihood loss. The ranking loss function uses this loss as the ranking basis to force the model to align the confidence score score of the key region with the classification performance, generating the ranking loss loss_key. When the model assigns a higher confidence score to a region with poor classification performance, this loss penalizes it, thereby prompting the model to learn that the ranking score of important regions is higher than that of unimportant regions.
[0180] S257: Add losspart1, losspart2, losspart3, losspart4, and loss_key to obtain L local .
[0181] S258: Transform the sorted key-region features fm1, fm2, and fm3 using a multi-layer perceptron, with the dimension becoming [16, 1024]. Perform regularization to ensure that the features after being processed by the key-region alignment module still retain the global semantic consistency. Calculate the distribution difference between the regularized features fm1, fm2, and fm3 and the global features f1, f2, and sp_feature through the regularization loss, and accumulate the obtained results to get L reg 。
[0182] S259: Calculate the loss L of the entire network.
[0183] L = L local +L global +L reg
[0184] In step S3, the following steps are included:
[0185] S31: Use the training set and test set obtained in S1 for training and testing.
[0186] S32: Set the parameters for training: the size of the input image is 448×448, the default number of training epochs is 300, and there are 16 samples in each epoch. The initial learning rate is 2e-3, and the learning rate for the backbone network is one-tenth of the initial learning rate. The backbone network uses ResNet-50 with pre-trained weights. Use the SGD optimizer, set the momentum to 0.9, and the weight decay to 5e-4. Learning rate scheduling: Use the cosine annealing scheduling strategy to dynamically adjust the learning rate within each epoch.
[0187] S33: In each round, verify once on the test set using the Top-1 accuracy, and save the state of the current model when the verification accuracy improves.
[0188] In step S4, after the model is trained multiple times, select the model with the highest accuracy for actual bird image recognition.
[0189] The embodiment of the present invention also provides a multi-scale feature extraction and region alignment device for fine-grained bird image recognition. The device includes:
[0190] An image acquisition module, used to acquire bird images; the bird images are the bird images to be recognized for fine-grained bird images;
[0191] An image preprocessing module, used to preprocess the bird images to generate standard bird images;
[0192] An image recognition module, configured to input the standard bird image into a pre-trained network model to predict the result of fine-grained bird image recognition.
[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
[0194] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium can include: ROM, RAM, disk or optical disc, etc.
[0195] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A fine-grained bird image recognition method for multi-scale feature extraction and region alignment, characterized in that It includes the following steps: S1: Collect bird images and construct a training and test data set; S2: Construct a fine-grained bird image recognition network for multi-scale feature extraction and region alignment; S3: After preprocessing the bird images, input them into the constructed model for training; S4: Use the fine-grained bird image recognition network for multi-scale feature extraction and region alignment to apply the trained model to actual bird image recognition.
2. The method for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to claim 1, characterized in that The specific steps in step S2 include the following steps: S21: Use a pre-trained backbone network to perform multi-scale feature extraction on a batch of input images I, and use a spatial attention mechanism to further enhance the features extracted by the backbone network; S22: Embed a branch activation structure and a self-attention feature representation mechanism in the backbone network for global feature aggregation to enhance the multi-scale perception ability and feature fusion ability of the backbone network; S23: Use a region proposal network based on dynamic anchor boxes to screen local key regions, and use the backbone network to extract local region features twice; S24: Align the local key regions to generate a stable key region feature representation; S25: Calculate the losses of each branch and the total loss.
3. The fine-grained bird image recognition method for multi-scale feature extraction and region alignment according to claim 2, characterized in that, The specific steps in step S21 include the following steps: S211: Return the intermediate features X={x1,...,x1,...,x2,...,x3,...,x4,...,x5,...,x6,...,x7,...,x8,...,x9,...,x10,...,x11,...,x12,...,x13,...,x14,...,x15,...,x16,...,x17,...,x18,...,x19,...,x219, m ,...,x n }, x m represents the feature map of the last convolution output of the mth main layer, n represents the number of main layers in the backbone network; S212: Obtain the feature maps Feature=(fs1, fs2) output by the last two convolutional layers of the last module of the last main layer of the backbone network; S213: The feature fs1 is used as an attention map to perform spatial attention weighting on the feature fs2 by means of inner product calculation. First, the spatial dimensions of fs1 and fs2 are flattened into one dimension, and the tensor multiplication of fs1 and fs2 in the spatial dimension is calculated using Einstein summation convention to obtain φ sp ; S214: Divide φ sp by the number of channels of the attention tensor and perform a square root operation, then perform L2 normalization on φ sp along the channel dimension, and adjust the finally obtained result back to the spatial dimension of the original feature fs2 to obtain the feature sp_feature.
4. The method for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to claim 2, wherein The specific steps in step S22 include the following steps: S221: Convolve the feature x with the shape of [B, C, H, W] output from the last layer of the backbone network through multiple parallel branches (each branch is a 1×1 convolutional layer), reducing the number of channels of the feature to 1 to generate multiple activation maps with the shape of [B, 1, H, W], Activations = {A1,..., A n ,..., A m ,..., A n}; A m = branch m (x n ) A m Represents the mth branch feature x n The activation maps generated by processing, n represents the total number of branches. These activation maps reflect the degree of attention of the model to different regions, similar to the mechanism of attention, indicating the importance of each spatial position. S222: Concatenate the activation maps together to obtain an activation map Activation with dimensions [B, n, H, W]. To achieve attention to important regions, unfold the activation map Activation in the spatial dimension, and its shape becomes [B, n, H×W]; S223: Use the Softmax function to normalize each spatial position of the activation map Activation to obtain a weighted activation map. Then perform a rearrangement operation on the activation map to make its shape become [B, H×W, n]; S224: Rearrange the original feature map x n to perform a rearrangement operation to change its shape to [B, C, H×W]. Weight x n using the activation map Activation, that is, multiply the feature values at each position of x n in the spatial dimension by the corresponding activation weights. The shape of the weighted feature N is [B, n, C]; N = x n × Softmax(Activation) S225: Change the shape of the feature N to [B, C, n] through a rearrangement operation; S226: Perform dimensionality reduction and non-linear mapping on the node feature N through a multi-layer perceptron to adjust it to a dimension suitable for input to the self-attention feature representation mechanism; S227: Regard the processed feature N as a sequence and input it into the self-attention feature representation mechanism for global interaction and aggregation; N attention = AttentionEncoder(MLP(N)).
5. The fine-grained bird image recognition method for multi-scale feature extraction and region alignment according to claim 2, wherein The specific steps in step S23 include the following steps: S231: Generate three feature maps with different resolutions by processing sp_feature through three convolutional layers to capture multi-scale target information; S232: The feature map sp_feature' generated from the m-th layer of convolution through a 1×1 convolution m predicts the original anchor box parameters P at each position m , and then converts it into a size s m and an aspect ratio r m . P m = Conv 1×1 (sp_feature' m ) s b,i,j,k,m = Softplus(p b,i,j,k,m ) + 10 -3 r b,i,j,k,m = Sigmoid(p b,i,j,k+K,m ) × 2 + 0.5 where P m is the original parameter tensor predicted by the m-th layer, with a shape of [B, 2K, H m , W m ; K is the number of anchor boxes at each position; s b,i,j,k,m is the size of the k-th anchor box of the feature generated by the m-th layer convolution of the b-th sample at the position (i, j); r b,i,j,k,m is the aspect ratio of the k-th anchor box of the m-th layer convolution of the b-th sample at the position (i, j); p b,i,j,k,m is the original value corresponding to the size in P m ; p b,i,j,k+K,m is the original value corresponding to the aspect ratio. S233: Generate a basic anchor box based on the predicted size s m and aspect ratio r m by combining the center coordinates of the feature map grid; S234: Predict the confidence and offset of each anchor box from the feature map sp_feature′ m ; where S m is the confidence of the feature map generated by the m-th layer of convolution, and after shape adjustment, it is [B, H m ×W m ×K]; D m is the offset of the feature map generated by the m-th layer of convolution, and after adjustment, the shape is [B, H m ×W m ×K, 4]. S235: Apply the offsets predicted by each layer of convolution to the base anchor boxes to generate the final bounding box B; S236: Screen out the topn regions through non-maximum suppression, and the confidence scores of these regions are stored in the list score; S237: Crop the region \(I\) corresponding to the bounding box \(B\) from the input image \(I\) part , and unify the size to \(224\times224\) by bilinear interpolation; I part = Interpolate(I[:, :, y′ 1,m :y′ 2,m , x′ 1,m :x′ 2,m , (224, 224)) where (x′ 1,m , y′ 1,m ) and (x′ 2,m , y2′ ,m ) represent the upper left corner coordinates and the lower right corner coordinates of the boundary predicted by the m-th layer convolution. S238: Feed the cropped image I part back into the backbone network to extract the features of the local key regions. When extracting features, use the same hierarchical feature extraction idea as the global branch, and layer by layer extract the local region features FP = {fp1,..., fp m ,..., fp n}, where the value range of m is from 1 to topn.
6. The method for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to claim 2, characterized in that The specific steps in step S24 include the following steps: S241: Generate random centers in the case of no initial samples in the initialization stage; Where C represents the cluster center matrix, with dimensions K×D, K represents the number of cluster centers, and D represents the feature dimension. S242: Create a backup feature center dictionary to back up each cluster center, ensuring that historical information is available for backtracking during the update process; C backup,k = C k , k = 0, 1, ..., K - 1 Among them, C backup,k represents the k-th backup clustering center, and C k represents the k-th initial clustering center. S243: Calculate the feature fp of each key area in the initial stage m The Euclidean distance to the cluster center k; D b,m,k = ||P b,m - C k ||², D ∈ R B×topn×K where D b,m,k represents the Euclidean distance from the m-th critical region to the k-th cluster center in batch b, and P m,n represents the feature vector of the m-th critical region in batch b. S244: Subsequently, take the index of the nearest center as the initial cluster label; Among which O b,m represents the index with the smallest Euclidean distance found among all the cluster centers for the m-th critical region in batch b. S245: Calculate the similarity between each pair of cluster centers using matrix multiplication, and form the relationship matrix A between the centers after normalization. center Calculate the pairwise similarity for all key regions within each image, and also perform normalization to obtain the relationship matrix A. S246: Construct a cost matrix M by multiplying the cluster center matrix and the critical region matrix and taking the negative value, and then use the linear programming method to find the optimal match O of batch b b , obtain the optimal permutation index of the critical regions of each image, so as to ensure that similar critical regions have a consistent order in different images. S247: When training the images of the next batch, each cluster center is updated according to the critical regions assigned to that center in the current batch. First, calculate the Euclidean distance between each critical region in the current batch and its corresponding cluster center. Use the exponential function to convert the distance into a weight, where the smaller the distance, the higher the weight. Normalize the weights within each sample to obtain w b,m , ensuring that the sum of the weights of all critical performance regions is 1. where w b,m represents the normalized weight of the m-th critical region in batch b; represents the cluster center to which the m-th critical region is assigned, O b,m is the corresponding index; Select all the key regions assigned to the cluster center k, and use the corresponding weight w b,m to multiply with the feature vector P of this key region b,m and then obtain the weighted feature sum of the cluster center k through summation. Then divide the weighted feature sum by the total weight of all the key regions assigned to the cluster center k to obtain the weighted average; where C′ k represents the weighted average center of the k-th cluster. To maintain the smoothness of updates, a smoothing factor α is introduced to control the degree of fusion between the new and old centers (C′ k is the new center calculated for this round, and C k is the old center obtained in the previous round). This mechanism of smooth update not only retains the information of the previous clustering centers but also can adjust the center position in a timely manner according to the data of the new batch. C k = αC k + (1 - α)C' k Among them, C k represents the updated cluster center. S248: Obtain the optimal matching order O of the current batch through the critical area alignment module b After that, reorder the features fp1, fp2, and fp3 to obtain fm1, fm2, and fm 3。 7. The method for fine-grained bird image recognition with multi-scale feature extraction and region alignment according to claim 2, wherein The specific steps in step S25 include the following steps: S251: The global feature branch processes the obtained global features f1, f2, and sp_feature through 3 different convolutional modules respectively. These 3 convolutional modules adopt a similar module combination, that is, first perform channel mapping by a 1×1 convolutional module, then extract local features by a 3×3 convolutional module, and finally compress the feature map into a 1×1 vector representation through an adaptive pooling layer. Then send them into different classifiers for classification to obtain results y1, y2, and y3. After concatenating f1, f2, and sp_feature, perform joint classification to obtain classification result y4. S252: Node feature N obtained by the auxiliary branch attention Similarly, the classifier is used for prediction to obtain the prediction result y5. S253: Calculate the losses for features y1, y2, y3, y4, and y5 respectively using smooth cross-entropy loss to obtain (loss1, loss2, loss3, loss4, loss5), and add these losses to get L global ; where p i represents a smoothing parameter to balance the contributions of predictions from each layer; z represents the true class label. S254: The processing of fp1, fp2, and fp3 is similar to that of global features. They are processed through a series of convolutional layers and adaptive pooling layers, and then sent to different classifiers for classification to obtain classification results yp1, yp2, and yp3. Use a concatenated classifier to concatenate fp1, fp2, and fp3 and perform joint classification to obtain classification result yp4; S255: Use the cross-entropy loss function to calculate the losses of these classification results to obtain (losspart1, losspart2, losspart3, losspart4); S256: Based on the predicted value yp4 of the critical region and the true label, calculate the classification loss of each region through the negative log-likelihood loss. The ranking loss function uses this loss as the ranking basis to force the model to align the confidence score score of the critical region with the classification performance, generating the ranking loss loss_key. When the model assigns a higher confidence score to a region with poor classification performance, this loss punishes it, thereby prompting the model to learn that the ranking score of important regions is higher than that of unimportant regions; S257: Add losspart1, losspart2, losspart3, losspart4 and loss_key to obtain L local ; S258: Regularize the sorted key-region features fm1, fm2, and fm3 to ensure that the features after being processed by the key-region alignment module still retain the global semantic consistency of the local features. Calculate the distribution differences between the regularized features fm1, fm2, and fm3 and the global features f1, f2, and sp_feature through the regularization loss, and accumulate the obtained results to get L reg ; S259: Calculate the loss L of the entire network; L = L local + L global + L reg。 8. A method and device for fine-grained bird image recognition with multi-scale feature extraction and region alignment, characterized in that The device includes: An image acquisition module for acquiring bird images; the bird images are bird images to be used for fine-grained bird image recognition; An image preprocessing module for preprocessing the bird images to generate standard bird images; An image recognition module for inputting the standard bird images into a pre-trained network model to predict the results of fine-grained bird image recognition.
Citation Information
Cited By
Power transformation knob equipment state identification method
CN120707906A