A network supervision fine-grained image recognition method and system based on deep learning

By obtaining images containing noise tags from the Internet, performing feature extraction and graph prototype construction, and using graph matching neural network model to optimize the recognition process, the problem of low fine-grained image recognition efficiency and accuracy in the prior art is solved, and efficient and accurate recognition results are achieved.

CN115496948BActive Publication Date: 2025-06-27GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211167812.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-27
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

The prior art has low efficiency and accuracy in fine-grained image recognition and relies on high-cost large-scale manual annotation data.

Method used

By obtaining input images containing noise tags from the Internet, performing feature extraction and instance diagram generation, constructing graph prototypes, and matching the instance diagram with graph prototype input diagram for neural network model training, optimize the model to improve recognition accuracy.

Benefits of technology

It significantly improves the efficiency and accuracy of fine-grained image recognition, reduces dependence on manual annotation data, and reduces data collection and annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496948B_ABST
    Figure CN115496948B_ABST
Patent Text Reader

Abstract

The present invention provides a network supervision fine-grained image recognition method and system based on deep learning. By performing feature processing on input images with noisy labels, instance graphs with noisy label features are obtained. Using the instance graphs with labels, graph prototypes are constructed for each category. The obtained instance graphs with noisy label features and the graph prototypes are used to train a pre-set graph matching neural network model, and the optimized graph matching neural network model is used for fine-grained image recognition. This method performs fine-grained image recognition based on deep learning. By introducing graph prototypes and performing contrastive learning with instance graphs with noisy label features, it can effectively correct noisy labels and eliminate outlier samples, significantly improving the efficiency and accuracy of fine-grained image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and more specifically, to a network supervised fine-grained image recognition method and system based on deep learning. Background Art

[0002] Fine-grained image recognition aims to identify subclasses of a given object category, such as different species of birds, as well as airplanes and cars, and has important scientific significance and application value in the fields of intelligent construction and the Internet. In recent years, with the continuous development of deep learning, great progress has been made in fine-grained image recognition.

[0003] Currently, most algorithms mainly use deep learning driven by high-quality data to achieve fine-grained image recognition, which largely depends on large-scale manually labeled data. The difficulty of collecting these datasets and the high cost of data annotation have become bottlenecks restricting their popularization and application.

[0004] At present, with the rapid development of the Internet, there is a large amount of weakly labeled data on the network that can be used to alleviate the dependence of current fine-grained image recognition algorithms on manual annotation, that is, the data obtained from network retrieval is used to train the neural network model. However, the data retrieved from the network contains a certain proportion of noisy labels, which will have an adverse impact on the training of the model. In addition, the inherent characteristics of small inter-class variance and large intra-class variance in fine-grained images further increase the difficulty of recognition.

[0005] The current prior art discloses a fine-grained image recognition algorithm based on distributed labels of inter-class similarity, including the following steps: using a backbone network to extract the feature representation of the input image; using a center loss module to calculate the center loss through the feature representation and update the class center; the classification loss module calculates the classification loss (such as cross-entropy loss) using the feature representation and the final label distribution, where the final label distribution is obtained by calculating the weighted sum of the one-hot label distribution and the distributed label distribution generated by the class center; the final objective loss function is obtained by weighted summing the center loss and the classification loss, and the entire model is optimized with this; the method in the prior art can alleviate the overfitting problem by reducing the confidence of model prediction, can effectively learn the discriminative features of fine-grained data, and can improve the accuracy of distinguishing different fine-grained category data to a certain extent; however, the method in the prior art mainly uses deep learning driven by high-quality data to distinguish subordinate categories, depends on large-scale manually labeled image data, has a high cost of data collection and annotation, is often time-consuming and laborious when performing fine-grained image recognition, and has problems of low efficiency and low accuracy. Summary of the Invention

[0006] To overcome the deficiencies of the above-mentioned prior art in terms of low efficiency and accuracy in fine-grained image recognition, the present invention provides a network-supervised fine-grained image recognition method and system based on deep learning, which can efficiently and accurately perform fine-grained recognition on images.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] A network-supervised fine-grained image recognition method based on deep learning includes the following steps:

[0009] S1: Obtain input images with noisy labels from the Internet;

[0010] S2: Extract features from the input images with noisy labels to obtain a region discriminant feature map and a global feature map;

[0011] S3: According to the obtained region discriminant feature map and global feature map, obtain instance maps with noisy label features;

[0012] S4: According to the obtained instance maps with noisy label features, construct graph prototypes for each category;

[0013] S5: Input the obtained instance maps with noisy label features and the graph prototypes into a pre-set graph matching neural network model for training to obtain an optimized graph matching neural network model;

[0014] S6: Obtain the image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized.

[0015] Preferably, in step S2, when extracting features from the input images with noisy labels to obtain a region discriminant feature map and a global feature map, the specific method is as follows:

[0016] Use a feature extractor to extract features from the input images with noisy labels to obtain a global feature map; pass the global feature map through a convolutional layer to obtain a globally averaged feature map after mean filtering; calculate the mean value of each position based on the number of channels of the globally averaged feature map after mean filtering to obtain a global mean feature map; search for the region with the maximum response value in the global mean feature map and locate the coordinates of the region with the maximum response value, and obtain a region discriminant feature map according to the coordinates of the region with the maximum response value.

[0017] Preferably, the specific method for searching for the region with the maximum response value in the global mean feature map and locating the coordinates of the region with the maximum response value is as follows:

[0018] Search for the region with the maximum response value in the global mean feature map and locate the coordinates of the region with the maximum response value according to the following formula:

[0019]

[0020]

[0021] Among them, represents the overall mean feature map, f‘ g represents the overall feature map after mean filtering, C represents the number of channels of the overall feature map after mean filtering, represents the row and column corresponding to the region for searching the maximum response value, and (i, j) represents the coordinates of the region of the maximum response value.

[0022] Preferably, in the step S3, according to the obtained region discrimination feature map and the overall feature map, an instance map containing noise label features is obtained. The specific method is as follows:

[0023] The obtained region discrimination feature map is transformed into the same dimension by using the method of bilinear interpolation to obtain a region feature map with the same dimension; the overall feature map and the region feature map with the same dimension are reduced in dimension by using the method of global average pooling to obtain a reduced overall feature map and a reduced region feature map; an instance map containing noise label features is obtained according to the reduced overall feature map and the reduced region feature map:

[0024] G ins =<V ins , E ins >

[0025] Among them, G ins represents the instance map containing noise label features, V ins represents the set of all feature points in the reduced overall feature map and the reduced region feature map, and E ins represents the adjacency matrix of the connections between feature points in the instance map containing noise label features.

[0026] Preferably, in the step S4, according to the obtained instance map containing noise label features, the specific method for constructing the graph prototype is as follows:

[0027] According to the obtained instance map containing noise label features, a graph prototype with the same structure as the instance map containing noise label features is constructed for each category, and the graph prototype is updated in a moving average manner:

[0028] G k =<V k , E k >

[0029]

[0030] Among them, G kRepresents the graph prototype of the k-th category constructed, V k represents the set of all feature points in the prototype of the k-th category, E k Represents the adjacency matrix connecting the feature points in the graph prototype of the kth category, G' k is the updated graph prototype, and m is the preset parameter.

[0031] Preferably, in step S5, the obtained instance graph containing noise label features and the graph prototype are input into a preset graph matching neural network model for training to obtain an optimized graph matching neural network model, and the specific method is:

[0032] The preset graph matching neural network model includes an intra-graph propagation layer, a graph aggregation layer, an inter-graph propagation layer and a graph matching layer, and obtaining the optimized graph matching neural network model includes the following steps;

[0033] S5.1: Obtained instance graph G containing noisy label features ins With the graph prototype G k Input the in-graph propagation layer to obtain the first feature matrix and the second feature matrix, and iteratively update the first feature matrix and the second feature matrix respectively through the graph convolution operation;

[0034] S5.2: Inputting the iteratively updated first feature matrix and the second feature matrix into the graph aggregation layer to combine features and obtain an aggregated feature vector;

[0035] S5.3: Input the aggregated feature vector into the inter-graph propagation layer for graph convolution operation, and iteratively update the aggregated feature vector to obtain the first feature expression f ins and the second characteristic expression Z k ;

[0036] S5.4: Express the first characteristic f ins and the second characteristic expression Z k Input graph matching layer to calculate similarity S k , according to the similarity S k Computing graph matching loss

[0037] S5.5: Correct the noise labels in the instance graph containing noise label features and remove outlier samples;

[0038] S5.6: Calculate classification cross entropy loss and total loss Based on total loss The graph matching neural network model is optimized to obtain an optimized graph matching neural network model.

[0039] Preferably, in the step S5.4, the first feature expression f ins and the second feature expression Z k are input into the graph matching layer to calculate the similarity S k , and the graph matching loss is calculated according to the similarity S k . Specifically: Specifically:

[0040] The first feature expression f ins and the second feature expression Z k are input into the graph matching layer for graph matching, and the similarity S k is calculated. Specifically:

[0041]

[0042] The graph matching layer sets a graph matching loss function, and the graph matching loss is calculated according to the similarity S k . The graph matching loss function is specifically:

[0043]

[0044]

[0045] Wherein, is the graph matching loss, y i represents the original label, k represents the category of the graph prototype, and K represents the total number of categories of the graph prototype.

[0046] Preferably, in the step S5.5, the noisy labels in the instance graph with noisy label features are corrected and the outlier samples are removed. The specific method is:

[0047] The graph inner propagation layer is provided with a classifier. The instance graph with noisy label features is input into the classifier to obtain the classifier distribution probability p i , and the graph matching distribution probability d i is calculated. The total probability q i is calculated according to the classifier distribution probability p i and the graph matching distribution probability d i . Specifically:

[0048] q i =αp i +(1 - α)d i

[0049]

[0050] Wherein, α is a preset parameter, and τ is a temperature coefficient;

[0051] According to the total probability q iCorrect the noisy labels in the instance graph containing noisy label features and remove the out-of-distribution samples OOD according to the total probability q and the preset threshold T, specifically as follows:

[0052]

[0053] Among them, is the pseudo-label, T is the preset threshold. When the maximum value of the total probability q i is greater than T, the category corresponding to the maximum value of the total probability q i is used as the pseudo-label; when the total probability q i is greater than the average probability of the category, the original label y i is used as the pseudo-label to correct the noisy labels in the instance graph containing noisy label features; in other cases, OOD is used as the pseudo-label, where OOD represents the out-of-distribution samples, to remove the out-of-distribution samples.

[0054] Preferably, in the step S5.6, calculate the classification cross-entropy loss and the total loss Optimize the graph matching neural network model according to the total loss to obtain the optimized graph matching neural network model. The specific method is as follows:

[0055] The graph inner propagation layer is provided with a classification cross-entropy loss function, specifically:

[0056]

[0057] Among them, is the classification cross-entropy loss, p ij is the classifier distribution probability of the i-th instance graph with noisy label features relative to the j-th category, is the pseudo-label of the i-th instance graph with noisy label features relative to the j-th category;

[0058] Construct a total loss function according to the classification cross-entropy loss function and the graph matching loss function. The total loss function is specifically:

[0059]

[0060] Among them, is the total loss, λ pro is the proportionality coefficient;

[0061] Optimize the graph matching neural network model according to the total loss to obtain the optimized graph matching neural network model.

[0062] The present invention also provides a network-supervised fine-grained image recognition system based on deep learning, which applies the above-mentioned network-supervised fine-grained image recognition method based on deep learning, and includes:

[0063] An image acquisition unit: used to acquire input images with noisy labels from the Internet;

[0064] A feature extraction unit: used to extract features from the input images with noisy labels to obtain a regional discriminative feature map and an overall feature map;

[0065] An instance graph generation unit: used to obtain instance graphs with noisy label features according to the obtained regional discriminative feature map and overall feature map;

[0066] A graph prototype construction unit: used to construct graph prototypes for each category according to the obtained instance graphs with noisy label features;

[0067] A graph matching unit: used to train the obtained instance graphs with noisy label features and the graph prototypes by inputting them into a preset graph matching neural network model to obtain an optimized graph matching neural network model;

[0068] An image recognition unit: used to obtain an image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized.

[0069] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0070] The present invention provides a network-supervised fine-grained image recognition method and system based on deep learning. The method processes the features of input images with noisy labels to obtain instance graphs with noisy label features, constructs a corresponding graph prototype for each category by using the instance graphs with noisy label features, trains the preset image matching neural network model with the obtained instance graphs with noisy label features and the graph prototypes and corrects the noisy labels, and uses the optimized image matching neural network model to perform fine-grained image recognition; the method performs network-supervised fine-grained image recognition based on deep learning, and can effectively correct the noisy labels by introducing contrast learning between the graph prototypes and the instance graphs with noisy label features, significantly improving the efficiency and accuracy of fine-grained image recognition. Description of the Drawings

[0071] Figure 1 It is a flowchart of a network-supervised fine-grained image recognition method provided for Embodiment 1.

[0072] Figure 2Schematic diagram of a network-supervised fine-grained image recognition method based on deep learning provided in Embodiment 2.

[0073] Figure 3 Structure diagram of a network-supervised fine-grained image recognition system based on deep learning provided in Embodiment 3.

[0074] 301 - Image acquisition unit, 302 - Feature extraction unit, 303 - Instance graph generation unit, 304 - Graph prototype construction unit, 305 - Graph matching unit, 306 - Image recognition unit. Detailed implementation manners

[0075] The attached drawings are only for illustrative purposes and should not be construed as limitations on this patent;

[0076] To better illustrate this embodiment, some components in the attached drawings are omitted, enlarged or reduced, which do not represent the sizes of actual products;

[0077] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0078] The technical solutions of the present invention will be further described below in conjunction with the attached drawings and embodiments.

[0079] Embodiment 1

[0080] As Figure 1 shown, this embodiment provides a network-supervised fine-grained image recognition method based on deep learning, including the following steps:

[0081] S1: Obtain input images with noisy labels from the Internet;

[0082] S2: Extract features from the input images with noisy labels to obtain region discriminant feature maps and overall feature maps;

[0083] S3: According to the obtained region discriminant feature maps and overall feature maps, obtain instance graphs with noisy label features;

[0084] S4: According to the obtained instance graphs with noisy label features, construct graph prototypes for each category;

[0085] S5: Input the obtained instance graphs with noisy label features and the graph prototypes into a pre-set graph matching neural network model for training to obtain an optimized graph matching neural network model;

[0086] S6: Obtain the image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized.

[0087] In the specific implementation process, first, obtain the input image with noisy labels through network retrieval. Then, use the CNN convolutional neural network to extract features from the input image with noisy labels to obtain the region discriminative feature map and the overall feature map. After that, obtain the instance map with the features of the noisy labels according to the obtained region discriminative feature map and the overall feature map. Then, construct a corresponding graph prototype for each category according to the instance map with the features of the noisy labels. Then, input the obtained instance map with the features of the noisy labels and the graph prototype into the pre-set graph matching neural network model for training, and calculate the graph matching loss and the classification cross-entropy loss to optimize the neural network, obtaining the optimized graph matching neural network model. Finally, use the optimized graph matching neural network model to identify the image to be recognized, and obtain the recognition result of the image to be recognized;

[0088] This method is based on deep learning for fine-grained image recognition. By introducing graph prototypes for contrastive learning with instance maps containing noisy label features, it can effectively correct noisy labels, significantly improving the efficiency and accuracy of fine-grained image recognition.

[0089] Embodiment 2

[0090] As Figure 2 shown, this embodiment provides a network-supervised fine-grained image recognition method based on deep learning, including the following steps:

[0091] S1: Obtain the input image with noisy labels from the Internet;

[0092] S2: Extract features from the input image with noisy labels to obtain the region discriminative feature map and the overall feature map. The specific method is as follows:

[0093] Use a feature extractor to extract features from the input image with noisy labels to obtain the overall feature map; pass the overall feature map through a convolutional layer to obtain the overall feature map after mean filtering; calculate the mean value of each position based on the number of channels of the overall feature map after mean filtering to obtain the overall mean feature map; search for the maximum response value region in the overall mean feature map, and locate the coordinates of the maximum response value region, and obtain the region discriminative feature map according to the coordinates of the maximum response value region;

[0094] The specific method for searching for the maximum response value region in the overall mean feature map and locating the coordinates of the maximum response value region is as follows:

[0095] Search for the maximum response value region in the overall mean feature map and locate the coordinates of the maximum response value region according to the following formula:

[0096]

[0097]

[0098] Among them, represents the overall mean feature map, f‘ g represents the overall feature map after mean filtering, C represents the number of channels of the overall feature map after mean filtering, represents the row and column corresponding to the region where the maximum response value is searched, and (i,j) represents the coordinates of the region with the maximum response value;

[0099] S3: According to the obtained region discrimination feature map and overall feature map, obtain an instance map containing noisy label features. The specific method is as follows:

[0100] Transform the obtained region discrimination feature map to the same dimension by using the method of bilinear interpolation to obtain a region feature map with the same dimension; use the method of global average pooling to reduce the dimension of the overall feature map and the region feature map with the same dimension to obtain the reduced-dimensional overall feature map and the reduced-dimensional region feature map; obtain an instance map containing noisy label features according to the reduced-dimensional overall feature map and the reduced-dimensional region feature map:

[0101] G ins = <V ins , E ins >

[0102] Among them, G ins represents an instance map containing noisy label features, V ins represents the set of all feature points in the reduced-dimensional overall feature map and the reduced-dimensional region feature map, and E ins represents the adjacency matrix of the connections between feature points in the instance map containing noisy label features;

[0103] S4: According to the obtained instance map containing noisy label features, construct a graph prototype for each category. The specific method is as follows:

[0104] According to the obtained instance map containing noisy label features, construct a graph prototype with the same structure as the instance map containing noisy label features for each category, and the graph prototype is updated in a moving average manner:

[0105] G k = <V k , E k >

[0106]

[0107] Among them, G k represents the graph prototype of the k-th category constructed, V k represents the set of all feature points in the graph prototype of the k-th category, and E kThe adjacency matrix representing the connections between feature points in the graph prototype of the k-th category, G' k is the updated graph prototype, and m is a preset parameter;

[0108] S5: Input the obtained instance graph with noisy label features and the graph prototype into a preset graph matching neural network model for training to obtain an optimized graph matching neural network model;

[0109] The preset graph matching neural network model includes an intra-graph propagation layer, a graph aggregation layer, an inter-graph propagation layer, and a graph matching layer. Obtaining the optimized graph matching neural network model includes the following steps;

[0110] S5.1: Input the obtained instance graph G ins with noisy label features and the graph prototype G k into the intra-graph propagation layer to obtain a first feature matrix and a second feature matrix, and respectively perform iterative updates on the first feature matrix and the second feature matrix through graph convolution operations, specifically:

[0111] Input the obtained instance graph G ins with noisy label features and the graph prototype G k into the intra-graph propagation layer, and reconstruct the set V of all feature points in the dimension-reduced overall feature map and the dimension-reduced regional feature map ins into a first feature matrix where n1 is the number of all feature points in the instance graph with noisy label features, and c1 is the dimension corresponding to each feature point in the instance graph with noisy label features;

[0112] Reconstruct the set V of all feature points in the graph prototype k into a second feature matrix where n2 is the number of all feature points in the graph prototype, and c2 is the dimension corresponding to each feature point in the graph prototype;

[0113] Perform graph convolution operations on the first feature matrix and the second feature matrix respectively, and iteratively update the first feature matrix and the second feature matrix, specifically:

[0114]

[0115]

[0116] where is the first feature matrix after the l-th iterative update, is the second feature matrix after the l-th iterative update, and are the parameters of the intra-graph propagation layer;

[0117] S5.2: Input the iteratively updated first feature matrix and second feature matrix into the graph aggregation layer for feature combination to obtain an aggregated feature vector, specifically:

[0118] Input the iteratively updated first feature matrix and second feature matrix into the image aggregation layer for feature combination to obtain an aggregated feature vector, specifically:

[0119]

[0120] Where, is the aggregated feature vector, is the updated first feature matrix, is the updated second feature matrix;

[0121] S5.3: Input the aggregated feature vector into the inter-graph propagation layer for graph convolution operation, and iteratively update the aggregated feature vector to obtain the first feature representation f ins and the second feature representation Z k , specifically:

[0122] Input the aggregated feature vector into the inter-graph propagation layer for graph convolution operation, and iteratively update the aggregated feature vector, specifically:

[0123]

[0124] Where, is the aggregated feature vector after the l-th iterative update, E cross is the adjacency matrix of the aggregated feature vector, and are the parameters of the inter-graph propagation layer;

[0125] Obtain the first feature representation f ins and the second feature representation Z k according to the aggregated feature vector after the l-th iterative update;

[0126] S5.4: Input the first feature representation f ins and the second feature representation Z k into the graph matching layer to calculate the similarity S k , and calculate the graph matching loss k according to the similarity S Specifically:

[0127] Input the first feature representation f ins and the second feature representation Z k into the graph matching layer for graph matching, and calculate the similarity S k , specifically:

[0128]

[0129] The graph matching layer sets a graph matching loss function, and calculates the graph matching loss according to the similarity S k The graph matching loss function is specifically:

[0130]

[0131]

[0132] where is the graph matching loss, y i represents the original label, k represents the category of the graph prototype, and K represents the total number of categories of the graph prototype;

[0133] S5.5: Correct the noise labels in the instance graph with noise label features and remove the outlier samples, specifically:

[0134] The graph inner propagation layer is provided with a classifier. Input the instance graph with noise label features into the classifier to obtain the classifier distribution probability p i and calculate the graph matching distribution probability d i According to the classifier distribution probability p i and the graph matching distribution probability d i calculate the total probability q i Specifically:

[0135] q i =αp i +(1 - α)d i

[0136]

[0137] where α is a preset parameter and τ is the temperature coefficient;

[0138] According to the total probability q i and the preset threshold T, correct the noise labels in the instance graph with noise label features and remove the outlier samples OOD, specifically:

[0139]

[0140] where is the pseudo label, T is the preset threshold. When the maximum value of the total probability q i is greater than T, the category corresponding to the maximum value of the total probability q i is used as the pseudo label; when the total probability q i is greater than the category average probability, the original label y iAs a pseudo-label, it is used to correct the noisy labels in the instance graph containing noisy label features; in other cases, the OOD (Out-of-Distribution) is used as a pseudo-label, where OOD represents outlier samples, to achieve the elimination of outlier samples.

[0141] S5.6: Calculate the categorical cross-entropy loss and the total loss According to the total loss optimize the graph matching neural network model to obtain an optimized graph matching neural network model, specifically:

[0142] The graph inner propagation layer is provided with a categorical cross-entropy loss function, specifically:

[0143]

[0144] Among them, is the categorical cross-entropy loss, p ij is the classifier distribution probability of the i-th instance graph with noisy label features relative to the j-th category, is the pseudo-label of the i-th instance graph with noisy label features relative to the j-th category;

[0145] Construct a total loss function according to the categorical cross-entropy loss function and the graph matching loss function, and the total loss function is specifically:

[0146]

[0147] Among them, is the total loss, λ pro is the proportionality coefficient;

[0148] According to the total loss optimize the graph matching neural network model to obtain an optimized graph matching neural network model;

[0149] S6: Obtain the image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized..

[0150] In the specific implementation process, first obtain the input image with noisy labels through network retrieval. The dataset used in this embodiment is WebFG-496, which consists of three sub-datasets, namely Web-Bird, Web-Aircraft, and Web-Car. The size of the input image with noisy labels is 448×448;

[0151] Afterwards, a convolutional neural network with ResNet50-varian as the backbone CNN is set up, and the feature extractor is used to extract features from the input image with noisy labels to obtain an overall feature map, where the dimension of the overall feature map is 14×14×2048; the overall feature map is passed through a convolutional layer to obtain the overall feature map after mean filtering; the mean value of each position is calculated based on the number of channels for the overall feature map after mean filtering to obtain the overall mean feature map;

[0152] Search for the region with the maximum response value in the overall mean feature map and locate the coordinates of the region with the maximum response value according to the following formula:

[0153]

[0154]

[0155] where, represents the overall mean feature map, f‘ g represents the overall feature map after mean filtering, C represents the number of channels of the overall feature map after mean filtering, represents the row and column corresponding to searching for the region with the maximum response value, and (i,j) represents the coordinates of the region with the maximum response value;

[0156] Intercept several local regions of different sizes in the overall feature map according to the coordinates of the obtained maximum response region. In this embodiment, three different area sizes S1, S2, S3 and three different aspect ratios A1, A2, A3, a total of 9 combinations are set to intercept the overall feature map, where the three different area sizes S1, S2, S3 are respectively one-half, one-third, and two-thirds of the area of the overall feature map, and the three different aspect ratio values A1, A2, A3 are respectively 1, 0.5, and 2;

[0157] Use the feature extractor to extract features from the intercepted local regions of different sizes to obtain region discriminant feature maps;

[0158] Construct an instance graph with noisy label features and a graph prototype corresponding to each category, and input the obtained instance graph with noisy label features and the graph prototype into the graph propagation layer GCN for graph convolution operations. In this embodiment, the output channel numbers are 1024 and 2048 respectively; aggregate the output instance graph with noisy label features and the graph prototype features, and obtain the first feature expression f ins and the second feature expression Z k ; Optimize the graph matching neural network model by calculating the graph matching loss and the classification cross-entropy loss respectively according to the first feature expression f ins and the second feature expression Z k ;

[0159] In this embodiment, α = 0.5, τ = 0.1, T = 0.75, λ pro = 1;

[0160] The images to be recognized are obtained from CUB200 - 2011, FGVC - Aircraft and Stanford Cars as verification data. After extracting the features of the images to be recognized, the optimized image matching neural network model is used to recognize the images to be recognized, and the recognition results of the images to be recognized are obtained;

[0161] As shown in the following table, it is a comparison chart of the recognition accuracy of fine - grained images by different methods:

[0162]

[0163] Table 1 - Comparison chart of the recognition accuracy of fine - grained images by different methods

[0164] Compared with the basic model, the method in this embodiment far exceeds various basic models in terms of performance on the three datasets. The backbone network used in this embodiment is ResNet - 50. Compared with the single ResNet - 50 model, the method in this embodiment has been greatly improved on the three datasets, and the average recognition accuracy has increased by 20.14%; for fair comparison, ResNet - 50 is uniformly used as the backbone network. It can be seen that when ResNet - 50 is used as the backbone network, the method in this embodiment achieves the highest average accuracy of 83.53%, and the accuracies on Web - Bird, Web - Aircraft and Web - Car are 76.62%, 85.79% and 82.09% respectively, which are 2.23%, 4.2% and 1.94% higher than the relatively advanced method Peer - learning; further using other models such as B - CNN as the backbone network, it can be seen from the comparison results that the method in this embodiment can be adapted to different backbone networks, thus obtaining a more obvious performance improvement in fine - grained image recognition; Figure 3 This method is based on deep learning for network - supervised recognition of fine - grained images. By introducing graph prototypes to perform contrastive learning with instance graphs containing noisy label features, it can effectively correct the noisy labels and significantly improve the efficiency and accuracy of fine - grained image recognition.

[0165]

[0166] Example 3

[0167] Figure 3 As Figure 3 shown, this embodiment provides a network - supervised fine - grained image recognition system based on deep learning, applying the network - supervised fine - grained image recognition method described in Example 1 or 2, including:

[0168] Image acquisition unit 301: used to obtain an input image with noisy labels from the Internet;

[0169] Feature extraction unit 302: used to extract features from the input image with noisy labels to obtain a region discrimination feature map and an overall feature map;

[0170] Instance graph generation unit 303: used to obtain an instance graph with the features of noisy labels according to the obtained region discrimination feature map and overall feature map;

[0171] Graph prototype construction unit 304: used to construct a graph prototype for each category according to the obtained instance graph with the features of noisy labels;

[0172] Graph matching unit 305: used to train the obtained instance graph with the features of noisy labels and the graph prototype by inputting them into a preset graph matching neural network model to obtain an optimized graph matching neural network model;

[0173] Image recognition unit 306: used to obtain an image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized;

[0174] In the specific implementation process, first, the image acquisition unit 301 is used to perform network retrieval to obtain an input image with noisy labels; then, the feature extraction unit 302 is used to extract features from the input image with noisy labels to obtain a region discrimination feature map and an overall feature map; the instance graph generation unit 303 is used to obtain an instance graph with the features of noisy labels according to the obtained region discrimination feature map and overall feature map; then, according to the obtained instance graph with the features of noisy labels, the graph prototype construction unit 304 is used to construct a graph prototype for each category; then, the graph matching unit 305 is used to train the obtained instance graph with the features of noisy labels and the graph prototype by inputting them into a preset graph matching neural network model to obtain an optimized graph matching neural network model; finally, the image recognition unit 306 obtains an image to be recognized, extracts the features of the image to be recognized, and then uses the optimized image matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized;

[0175] This system is based on deep learning for fine-grained image recognition. By introducing graph prototypes and comparing and learning with instance graphs with the features of noisy labels, it can effectively correct the noisy labels and significantly improve the efficiency and accuracy of fine-grained image recognition.

[0176] The same or similar reference numerals correspond to the same or similar components;

[0177] The terms used to describe the positional relationship in the attached drawings are for illustrative purposes only and should not be construed as a limitation of this patent;

[0178] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A network supervised fine-grained image recognition method based on deep learning, characterized in that, It includes the following steps: S1: Obtain an input image with noisy labels from the Internet; S2: Extract features from the input image with noisy labels to obtain a regional discriminant feature map and a global feature map; S3: According to the obtained regional discriminant feature map and global feature map, obtain an instance map with noisy label features; S4: According to the obtained instance map with noisy label features, construct a graph prototype for each category; S5: Input the obtained instance map with noisy label features and the graph prototype into a pre-set graph matching neural network model for training to obtain an optimized graph matching neural network model; The pre-set graph matching neural network model includes an intra-graph propagation layer, a graph aggregation layer, an inter-graph propagation layer, and a graph matching layer. Obtaining an optimized graph matching neural network model includes the following steps; S5.1: Input the obtained instance map with noisy label features and the graph prototype into the intra-graph propagation layer to obtain a first feature matrix and a second feature matrix, and respectively perform iterative updates on the first feature matrix and the second feature matrix through graph convolution operations; S5.2: Input the iteratively updated first feature matrix and second feature matrix into the graph aggregation layer for feature combination to obtain an aggregated feature vector; S5.3: Input the aggregated feature vector into the inter-graph propagation layer for graph convolution operation, and iteratively update the aggregated feature vector to obtain the first feature representation and the second feature representation ; S5.4: Input the first feature representation and the second feature representation into the graph matching layer to calculate the similarity , and calculate the graph matching loss based on the similarity ; S5.5: Correct the noisy labels in the instance map with noisy label features and remove outlier samples; S5.6: Calculate the categorical cross-entropy loss and the total loss , and optimize the graph matching neural network model according to the total loss to obtain an optimized graph matching neural network model; S6: Obtain an image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized.

2. The method for network-supervised fine-grained image recognition based on deep learning according to claim 1, wherein In the step S2, when extracting features from the input image with noisy labels to obtain a regional discriminant feature map and a global feature map, the specific method is as follows: Use a feature extractor to extract features from the input image with noisy labels to obtain a global feature map; pass the global feature map through a convolutional layer to obtain a globally averaged feature map after filtering; calculate the mean value of each position based on the number of channels of the globally averaged feature map after filtering to obtain a global mean feature map; search for the region with the maximum response value in the global mean feature map and locate the coordinates of the region with the maximum response value, and obtain a regional discriminant feature map according to the coordinates of the region with the maximum response value.

3. A fine-grained image recognition method based on deep learning network supervision according to claim 2, characterized in that, The specific method for searching for the region with the maximum response value in the global mean feature map and locating the coordinates of the region with the maximum response value is as follows: Search for the region with the maximum response value in the global mean feature map and locate the coordinates of the region with the maximum response value according to the following formula: Among them, represents the overall mean feature map, represents the overall feature map after mean filtering, C represents the number of channels of the overall feature map after mean filtering, represents the rows and columns corresponding to the region of the maximum response value search, represents the coordinates of the maximum response value region.

4. A fine-grained image recognition method based on deep learning network supervision according to claim 3, characterized in that, In the step S3, according to the obtained regional discriminant feature map and global feature map, to obtain an instance map with noisy label features, the specific method is as follows: Transform the obtained regional discriminant feature map to the same dimension by using the method of bilinear interpolation to obtain a regional feature map with the same dimension; use the method of global average pooling to reduce the dimension of the global feature map and the regional feature map with the same dimension to obtain a reduced-dimensional global feature map and a reduced-dimensional regional feature map; obtain an instance map with noisy label features according to the reduced-dimensional global feature map and the reduced-dimensional regional feature map: Among them, represents an instance graph containing noisy label features, represents the set of all feature points in the overall feature graph after dimensionality reduction and the regional feature graph after dimensionality reduction, represents the adjacency matrix of the connections between feature points in the instance graph containing noisy label features.

5. A network supervised fine-grained image recognition method based on deep learning according to claim 4, characterized in that, In step S4, the specific method for constructing a graph prototype according to the obtained instance graph with noisy label features is as follows: According to the obtained instance graph with noisy label features, a graph prototype with the same structure as the instance graph with noisy label features is constructed for each category, and the graph prototype is updated by means of moving average: Among them, represents the graph prototype of the k-th category constructed, represents the set of all feature points in the graph prototype of the k-th category, represents the adjacency matrix of the connections between feature points in the graph prototype of the k-th category, is the updated graph prototype, and m is a preset parameter.

6. A network supervised fine-grained image recognition method based on deep learning according to claim 5, characterized in that In the step S5.4, the first feature expression and the second feature expression are input into the graph matching layer to calculate the similarity , and the graph matching loss is calculated according to the similarity , specifically as follows: Input the first feature representation and the second feature representation into the graph matching layer for graph matching and calculate the similarity , specifically as follows: The graph matching layer sets a graph matching loss function and calculates the graph matching loss according to the similarity The specific form of the graph matching loss function is as follows: Among them, is the graph matching loss, represents the original label, k represents the category of the graph prototype, and K represents the total number of categories of the graph prototypes.

7. A fine-grained image recognition method based on deep learning network supervision according to claim 6, characterized in that In step S5.5, the specific method for correcting the noisy labels in the instance graph with noisy label features and removing the outlier samples is as follows: The in-graph propagation layer is provided with a classifier. The instance graph containing the noisy label features is input into the classifier to obtain the classifier distribution probability , calculate the graph matching distribution probability , according to the classifier distribution probability and the graph matching distribution probability calculate the total probability , specifically: Among them, is a preset parameter, is the temperature coefficient; According to the total probability and a preset threshold T, correct the noisy labels in the instance graph containing noisy label features and remove the out-of-distribution samples OOD, specifically as follows: Among them, is a pseudo-label, T is a preset threshold. When the maximum value of the total probability is greater than T, the category corresponding to the maximum value of the total probability is used as the pseudo-label; when the total probability is greater than the average probability of the category, the original label is used as the pseudo-label to correct the noisy labels in the instance graph with noisy label features; in other cases, OOD (out-of-distribution samples) is used as the pseudo-label to remove out-of-distribution samples.

8. A fine-grained image recognition method based on deep learning network supervision according to claim 7, characterized in that In the step S5.6, calculate the categorical cross-entropy loss and the total loss . According to the total loss , optimize the graph matching neural network model to obtain an optimized graph matching neural network model. The specific method is as follows: The classification cross-entropy loss function is set in the intra-graph propagation layer, specifically: Among them, is the categorical cross-entropy loss, is the classifier distribution probability of the i-th instance graph with noisy label features relative to the j-th category, is the pseudo-label of the i-th instance graph with noisy label features relative to the j-th category; A total loss function is constructed according to the classification cross-entropy loss function and the graph matching loss function. The total loss function is specifically: Among them, is the total loss, is the proportionality coefficient; According to the total loss Optimize the graph matching neural network model to obtain an optimized graph matching neural network model.

9. A network-supervised fine-grained image recognition system based on deep learning, which applies a network-supervised fine-grained image recognition method described in any one of claims 1-8, and is characterized in that, including: Image acquisition unit: used to acquire input images with noisy labels from the Internet; Feature extraction unit: used to extract features from the input images with noisy labels to obtain a regional discriminative feature map and an overall feature map; Instance graph generation unit: used to obtain an instance graph with noisy label features according to the obtained regional discriminative feature map and overall feature map; Graph prototype construction unit: used to construct a graph prototype for each category according to the obtained instance graph with noisy label features; Graph matching unit: used to input the obtained instance graph with noisy label features and the graph prototype into a pre-set graph matching neural network model for training to obtain an optimized graph matching neural network model; Image recognition unit: used to acquire the image to be recognized, extract the features of the image to be recognized, and then use the optimized graph matching neural network model to recognize the image to be recognized to obtain the recognition result of the image to be recognized.

Citation Information

Patent Citations

  • Small sample image recognition method based on deep learning

    CN109800811A

  • Image fine-grained classification method, system and equipment

    CN113392875A