Fine-grained image retrieval method and system based on multi-branch multi-loss learning

Through a hash network with multi-branch and multi-loss learning, the problem of insufficient feature differential capture in fine-grained image retrieval is solved, and more efficient image retrieval accuracy and efficiency are achieved.

CN120372038APending Publication Date: 2025-07-25SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499612.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25

Smart Images

  • Figure CN120372038A_ABST
    Figure CN120372038A_ABST
Patent Text Reader

Abstract

The invention provides a fine-grained image retrieval method and system based on multi-branch and multi-loss learning, and relates to the field of image processing.The method comprises the steps that a Hash network based on multi-branch and multi-loss learning is established and at least comprises a basic network, a feature filtering module and a Hash network, the basic network is used for extracting a feature map of an input image, the feature filtering module is used for removing noise features in the channel and spatial dimensions of the feature map and generating a filtered feature map, and the Hash network is used for generating a Hash code based on the filtered feature map; according to the multi-stage loss function, training a multi-branch multi-loss learning-based Hash network; generating a Hash code of the sample image through the trained multi-branch multi-loss learning Hash network; generating a Hash code of a query image through the trained multi-branch multi-loss learning Hash network; and generating an image retrieval result based on the hash code of the sample image and the hash code of the query image, so that the method has the advantage of improving the accuracy of fine-grained image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a fine-grained image retrieval method and system based on multi-branch multi-loss learning. Background Art

[0002] In recent years, deep hashing learning methods have achieved remarkable success in large-scale image retrieval tasks due to their high efficiency and low storage requirements. However, when directly applied to fine-grained image retrieval (FGIR), their performance significantly degrades.

[0003] The FGIR task requires differentiating highly similar subclasses (such as different breeds of birds, vehicle models, etc.), which have significant intra-class differences (such as pose, lighting, occlusion, etc.) in visual features, while the inter-class differences are very subtle. Existing hashing methods usually rely on global features (such as the output of the fully connected layer of a CNN (Convolutional Neural Network)), making it difficult to capture these subtle local differences. For example, with the failure of global features, different poses of the same bird species may lead to significant differences in global features, while the subtle differences between different bird species are ignored; with the lack of local features, the hashing method does not fully explore the key local regions in the image (such as the beak and wing texture of a bird), resulting in the generated binary codes lacking discriminability. In FGIR, the semantic similarity of images (such as "the same bird species") should be consistent with the generated hash codes. However, during the optimization process, many hashing methods only focus on the quantization error. When mapping continuous features to binary codes, they minimize the quantization error (such as minimizing the difference between real-valued features and binary codes), while ignoring the semantic similarity. At the same time, during the feature embedding or hash code generation process, the semantic relationship is not explicitly modeled, resulting in large differences in the hash codes of similar images. Consequently, in the retrieval results, semantically similar images are wrongly assigned to different hash buckets, further reducing the retrieval performance.

[0004] Therefore, there is a need to provide a fine-grained image retrieval method and system based on multi-branch multi-loss learning to improve the accuracy of fine-grained image retrieval. Summary of the Invention

[0005] The present invention provides a fine-grained image retrieval method based on multi-branch multi-loss learning, including: establishing a hash network based on multi-branch multi-loss learning, where the hash network based on multi-branch multi-loss learning at least includes a basic network, a feature filtering module, and a hash network. The basic network is used to extract the feature map of the input image. The feature filtering module is used to remove noise features in the channel and spatial dimensions of the feature map to generate a filtered feature map. The hash network is used to generate a hash code based on the filtered feature map; training the hash network based on multi-branch multi-loss learning according to a multi-stage loss function; generating a hash code of a sample image through the trained hash network based on multi-branch multi-loss learning; obtaining a query image; generating a hash code of the query image through the trained hash network based on multi-branch multi-loss learning; and generating an image retrieval result based on the hash code of the sample image and the hash code of the query image.

[0006] Further, the basic network includes a feature extraction unit, a double-selection saliency region erasing unit, and a double-filtering target localization unit. The feature extraction unit includes a first feature extraction branch, a second feature extraction branch, and a third feature extraction branch. Among them, the first feature extraction branch is used to perform preliminary feature extraction on the image to obtain a preliminary feature map. The double-selection saliency region erasing unit is used to generate a region-erased image of the input image based on the preliminary feature map. The second feature extraction branch is used to extract the feature map of the region-erased image. The double-filtering target localization unit is used to generate a target localization image of the input image based on the preliminary feature map. The third feature extraction branch is used to extract the feature map of the target localization image. Among them, the feature map of the input image includes the preliminary feature map, the feature map of the region-erased image, and the feature map of the target localization image.

[0007] Further, the double-selection saliency region erasing unit includes a channel filtering group and a spatial erasing group. The channel filtering group is used to calculate the proportion of non-zero pixels in each channel of the preliminary feature map, calculate the average value of the non-zero pixel proportions of all channels, delete the channels with proportions lower than the average value to obtain a channel-filtered feature map. The spatial erasing group is used to perform channel aggregation on the channel-filtered feature map, adjust the preliminary feature map to the same size as the input image through bilinear interpolation, erase the region with the highest response based on a threshold, and multiply it element by element with the input image to generate a region-erased image of the input image.

[0008] Further, the dual-filter target localization unit includes a spatial filter group and a noise background filter group; the spatial filter group is used to perform channel aggregation on the preliminary feature map, perform an average operation on each channel of the feature map after channel aggregation in the spatial dimension, remove features smaller than the average value, and generate a spatially filtered feature map; the noise background filter group is used to filter the noise features of the spatially filtered feature map by the largest connected component method, detect the position of the key area, generate a region mask of the same size as the input image based on the position of the key area using bilinear interpolation, and multiply the region mask and the input image element by element to generate the target localization image of the input image.

[0009] Further, the feature filtering module includes a first feature filtering branch, a second feature filtering branch, and a third feature filtering branch; the first feature filtering branch is used to perform feature filtering on the preliminary feature map to generate a filtered first branch feature map; the second feature filtering branch is used to perform feature filtering on the region erasure image of the input image to generate a filtered second branch feature map; the third feature filtering branch is used to perform feature filtering on the target localization image of the input image to generate a filtered third branch feature map, where the filtered feature map includes the filtered first branch feature map, the filtered second branch feature map, and the filtered third branch feature map.

[0010] Further, the first feature filtering branch includes a channel feature filtering group and a spatial feature filtering group; the channel feature filtering group is used to calculate the proportion of non-zero values in each channel of the preliminary feature map, and delete channels smaller than the proportion threshold to generate a channel feature filtered feature map; the spatial feature filtering group is used to perform channel aggregation processing on the channel feature filtered feature map to generate an aggregated feature map, and generate a filtered first branch feature map based on the feature value and the global average value at each spatial position of the aggregated feature map.

[0011] Further, the hash network includes at least a third fully connected layer and a hash layer: the third fully connected layer is used to perform operations on the filtered third branch feature map and output the classification features of the filtered third branch feature map; the hash layer is used to generate a hash code based on the classification features and classification weights of the filtered third branch feature map.

[0012] Furthermore, the hash network further includes a first fully-connected layer and a second fully-connected layer; the first fully-connected layer is used to perform operations on the filtered first-branch feature map and output the classification features of the filtered first-branch feature map; the second fully-connected layer is used to perform operations on the filtered second-branch feature map and output the classification features of the filtered second-branch feature map; the multi-stage loss function includes a saliency loss, a classification loss, an adversarial loss, a target loss, and a hash loss calculated based on the classification features of the filtered first-branch feature map, the classification features of the filtered second-branch feature map, and the classification features of the filtered third-branch feature map.

[0013] Furthermore, the saliency loss is calculated based on the classification features of the filtered second-branch feature map; the classification loss is calculated based on the classification features of the filtered first-branch feature map; the adversarial loss is calculated based on the classification features of the filtered first-branch feature map and the classification features of the filtered third-branch feature map; the target loss is calculated based on the filtered third-branch feature map and the classification features of the filtered third-branch feature map; the hash loss is calculated based on the hash code.

[0014] The present invention provides a fine-grained image retrieval system based on multi-branch multi-loss learning for the above-mentioned fine-grained image retrieval method based on multi-branch multi-loss learning, including: a model establishment module for establishing a hash network based on multi-branch multi-loss learning, where the hash network based on multi-branch multi-loss learning at least includes a basic network, a feature filtering module, and a hash network, the basic network is used to extract the feature map of the input image, the feature filtering module is used to remove noise features in the channel and spatial dimensions of the feature map and generate a filtered feature map, and the hash network is used to generate a hash code based on the filtered feature map; a model training module for training the hash network based on multi-branch multi-loss learning according to the multi-stage loss function; a hash code generation module for generating the hash code of the sample image through the trained hash network based on multi-branch multi-loss learning; an image acquisition module for acquiring a query image; the hash code generation module is further used to generate the hash code of the query image through the trained hash network based on multi-branch multi-loss learning; an image retrieval module for generating an image retrieval result based on the hash code of the sample image and the hash code of the query image.

[0015] Compared with the prior art, the fine-grained image retrieval method and system based on multi-branch multi-loss learning provided by the present invention at least have the following beneficial effects:

[0016] Through the collaborative work of multiple branches such as the basic network and the feature filtering module, it is possible to more comprehensively capture the fine-grained features in the image, including local details and global structures, and effectively address the challenges of large intra-class variations and small inter-class variations in fine-grained image retrieval.

[0017] The feature filtering module removes noise features in the channel and spatial dimensions of the feature map, which helps to highlight key features and improve the purity and discriminability of feature representation.

[0018] By erasing the salient regions, it prompts the multi-branch multi-loss learning hashing network to learn more compact and discriminative binary codes, which helps to reduce intra-class similarity and increase inter-class difference.

[0019] Design a dedicated multi-stage loss function that can be optimized according to the characteristics of the fine-grained image retrieval task, effectively improving the discrimination ability of the binary codes, thereby enhancing the retrieval accuracy. Training with a multi-stage loss function can gradually guide the network to learn better feature representations and binary code generation strategies, which helps the network converge faster and achieve better performance.

[0020] By generating the hash codes of the sample images and query images through the trained network and performing similarity measurement based on these hash codes, the image retrieval results can be quickly generated, significantly improving the retrieval efficiency. By more accurately capturing fine-grained features and generating discriminative binary codes, the retrieval accuracy is also significantly improved. Brief Description of the Drawings

[0021] This specification will be further described by way of exemplary embodiments, which will be described in detail through the accompanying drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0022] Figure 1 is a schematic flowchart of a fine-grained image retrieval method based on multi-branch multi-loss learning according to some embodiments of this specification;

[0023] Figure 2 is a schematic structural diagram of a multi-branch multi-loss learning hashing network according to some embodiments of this specification;

[0024] Figure 3 is a schematic structural diagram of a dual-selection salient region erasing unit according to some embodiments of this specification;

[0025] Figure 4 is a schematic structural diagram of a dual-filtering target localization unit according to some embodiments of this specification;

[0026] Figure 5 is a schematic module diagram of a fine-grained image retrieval system based on multi-branch multi-loss learning according to some embodiments of this specification. Detailed Description of the Embodiments

[0027] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the figures represent the same structure or operation.

[0028] Figure 1 is a schematic flowchart of a fine-grained image retrieval method based on multi-branch multi-loss learning according to some embodiments of this specification, as Figure 1 shown, the fine-grained image retrieval method based on multi-branch multi-loss learning may include the following steps.

[0029] Step 110, establish a hash network based on multi-branch multi-loss learning.

[0030] Figure 2 is a schematic structural diagram of a hash network based on multi-branch multi-loss learning according to some embodiments of this specification, as Figure 2 shown, the hash network based on multi-branch multi-loss learning at least includes a basic network, a feature filtering module, and a hash network. The basic network is used to extract the feature map of the input image. The feature filtering module is used to remove the noise features in the channel and spatial dimensions of the feature map to generate a filtered feature map. The hash network is used to generate a hash code based on the filtered feature map.

[0031] Specifically, the basic network component initially extracts the feature map of the input image by adopting ResNet18 and the Dual Select Significant Region Erasure (DSSRE) and Dual Filtering Object Localization (DFOL) units, so as to provide prior information with multiple branches for subsequent components. The feature filtering module removes the noise features in the channel and spatial dimensions of the feature map through the Channel Filtering (CF) group and the Spatial Filtering (SF) group, so as to extract discriminative feature information. The hash network takes the filtered feature map as the input and generates a binary hash code.

[0032] The definitions of each parameter are shown in Table 1.

[0033] Table 1

[0034]

[0035]

[0036] asFigure 2 As shown, preferably, the basic network includes a feature extraction unit, a dual-selection saliency region erasing unit, and a dual-filtering target localization unit.

[0037] The feature extraction unit includes a first feature extraction branch, a second feature extraction branch, and a third feature extraction branch. Among them, the first feature extraction branch is used to perform preliminary feature extraction on the image to obtain a preliminary feature map.

[0038] The dual-selection saliency region erasing unit is used to generate a region-erased image of the input image based on the preliminary feature map.

[0039] The second feature extraction branch is used to extract the feature map of the region-erased image.

[0040] The dual-filtering target localization unit is used to generate a target localization image of the input image based on the preliminary feature map.

[0041] The third feature extraction branch is used to extract the feature map of the target localization image. Among them, the feature map of the input image includes the preliminary feature map, the feature map of the region-erased image, and the feature map of the target localization image.

[0042] Specifically, to solve the problem of low retrieval accuracy caused by large intra-class differences and small inter-class differences in fine-grained images, the first feature extraction branch uses ResNet to perform preliminary feature extraction on the input image x i to obtain a preliminary feature map E. Through the dual-selection saliency region erasing unit, the region-erased image x i of the input image x can be obtained. i e At the same time, the target localization image x i of the input image x is extracted through the dual-filtering target localization unit. i o . Both the second feature extraction branch and the third feature extraction branch use ResNet.

[0043] Figure 3 is a schematic structural diagram of the dual-selection saliency region erasing unit shown according to some embodiments of the present specification. As Figure 3 shown, preferably, the dual-selection saliency region erasing unit (DSSRE) includes a channel filtering group and a spatial erasing group.

[0044] The channel filtering group is used to calculate the proportion of non-zero pixels in each channel of the preliminary feature map, calculate the average value of the non-zero pixel proportions of all channels, and delete the channels lower than the average value to obtain a channel-filtered feature map.

[0045] The spatial erasure group is used to perform channel aggregation on the feature map after channel filtering, and adjust the preliminary feature map to the same size as the input image through bilinear interpolation, erase the region with the highest response based on a threshold, and multiply it element-wise with the input image to generate the region-erased image of the input image.

[0046] Specifically, in the fine-grained image retrieval task, most previous methods only extract discriminative features in the feature map through a single random mask or attention mask method, but ignore the noise filtering on the channels, resulting in inaccurate discriminative features being extracted. To solve this problem, the dual-selection saliency region erasure unit completes feature extraction through two steps: channel filtering and spatial erasure.

[0047] The purpose of the channel filtering operation is to remove the noise feature channels and highlight the feature channels with strong discrimination ability. Usually, the features extracted by the basic network contain a large number of noise channels, which cause serious interference to object discrimination. Especially in fine-grained tasks, more refined discriminative features are required. For example, when the pixel values in a certain channel are large and mostly non-zero, visually, white almost occupies the entire channel area, and it can be inferred that the feature map of such channels is not conducive to object region localization. Therefore, these channels can be deleted; on the contrary, for the channel feature map with a relatively small proportion of the white area, it contains more discriminative information and is more conducive to fine-grained tasks. Therefore, these channel feature maps should be retained. Specifically, calculate the proportion of non-zero pixels in each channel, and then calculate the average value of the non-zero pixel proportions of all channels, and delete the channels below the average value to obtain the feature map after channel filtering. The calculation formula of the channel filtering (CF) group is:

[0048]

[0049] where Res(x i ) ∈ R c×w×h represents the feature map extracted by ResNet, and the channel filtering mask cm i of x i is expressed as:

[0050]

[0051] where represents the proportion of non-zero elements in the nth channel, and the average value of the non-zero channel proportions is:

[0052]

[0053] The spatial erasing operation erases the most discriminative part of the original image by setting a threshold, thereby generating a compact binary code. This operation destroys the texture of the original image, forcing the network to focus on the shape of the target and exploit other neglected but discriminative features. Specifically, after taking the output feature map of the channel filtering module as input, the spatial erasing (SE) group first performs channel aggregation and resizes the feature map to the same size as the original image by bilinear interpolation. Then, the area with the highest response is erased by setting a threshold and multiplied element-wise with the original image to obtain the output feature map. The output of the bilinear interpolation κ i It is expressed as:

[0054]

[0055] Spatial erasure mask sem i It is expressed as follows:

[0056]

[0057] Among them, α represents the set threshold, which is used to determine the size of the area to be erased; max(·) represents the maximum value in the feature map. Finally, the output of the dual selection salient region erasure unit (DSSRE) is calculated as follows:

[0058]

[0059] Figure 4 is a schematic diagram of the structure of a dual filtering target positioning unit according to some embodiments of this specification, such as Figure 4 As shown, preferably, the dual filtering target positioning unit includes a spatial filtering group and a noise background filtering group.

[0060] The spatial filtering group is used to perform channel aggregation on the preliminary feature map, perform an average operation on each channel of the feature map after channel aggregation in the spatial dimension, remove features smaller than the average value, and generate a feature map after spatial filtering;

[0061] The noise background filter group is used to filter the noise features of the feature map after spatial filtering through the maximum connected component method, and detect the position of the key area. Based on the position of the key area, bilinear interpolation is used to generate a region mask of the same size as the input image, and the region mask is multiplied element by element with the input image to generate a target positioning image of the input image.

[0062] Specifically, compared with coarse-grained images, fine-grained image retrieval tasks are difficult to accurately locate key regions and extract effective features from the detected key regions due to large intra-class differences and small inter-class differences. To address this issue, the Dual Filter Object Localization Unit (DFOL) removes features smaller than the average through channel aggregation and spatial averaging, thereby obtaining a spatially filtered feature map and providing the most useful prior information for the subsequent network.

[0063] Taking the feature map Res(x i ) ∈ R c×w×h extracted by ResNet as input, the calculation formula for channel aggregation is:

[0064]

[0065] The expression of the spatial filtering group is as follows:

[0066]

[0067] The noise background filtering operation filters out the noise features in the space through the largest connected component method and detects the position of the key region. Subsequently, a region mask of the same size as the original image is obtained using bilinear interpolation. Then, the mask is multiplied element-wise with the original image to generate the output of the Dual Filter Object Localization Unit (DFOL) Its calculation formula is as follows:

[0068]

[0069] Preferably, the feature filtering module includes a first feature filtering branch, a second feature filtering branch, and a third feature filtering branch.

[0070] The first feature filtering branch is used to perform feature filtering on the preliminary feature map to generate a filtered first branch feature map;

[0071] The second feature filtering branch is used to perform feature filtering on the region-erased image of the input image to generate a filtered second branch feature map;

[0072] The third feature filtering branch is used to perform feature filtering on the object localization image of the input image to generate a filtered third branch feature map, where the filtered feature map includes the filtered first branch feature map, the filtered second branch feature map, and the filtered third branch feature map.

[0073] Preferably, the first feature filtering branch includes a channel feature filtering group and a spatial feature filtering group.

[0074] The channel feature filtering group is used to calculate the proportion of non-zero values in each channel of the preliminary feature map and delete the channels smaller than the proportion threshold to generate a channel feature-filtered feature map;

[0075] The spatial feature filtering group is used to perform channel aggregation processing on the feature map after channel feature filtering to generate an aggregated feature map, and based on the feature values and global average values at each spatial position of the aggregated feature map, a filtered first-branch feature map is generated.

[0076] Specifically, channel feature filtering aims to remove noisy channels and feature channels lacking discriminative ability. Since the operations of the first feature filtering branch, the second feature filtering branch, and the third feature filtering branch are the same, here the filtering process is described in detail taking the first feature filtering branch as an example: Considering a threshold υ (usually obtained through extensive experimental research), the first feature filtering branch takes the feature map E∈R c×w×h generated by ResNet as input, calculates the proportion of non-zero values in each channel, and deletes the channels with values less than the threshold υ.

[0077] Spatial feature filtering (SFF) aims to further remove spatial noise in the feature map and extract more discriminative features. First, channel aggregation processing is performed on the feature map of CFF:

[0078]

[0079] The output of the spatial feature filtering group is expressed as:

[0080]

[0081] where u is a specific value ranging from 0 to 1 (see the implementation settings in the experimental part for details). Generally speaking, the output of the feature filtering component is:

[0082] f i = SFF(CFF(Res(x i )))

[0083] Similarly, the second feature filtering branch and the third feature filtering branch respectively take and as input for feature filtering and output and

[0084] Preferably, the hash network includes at least a third fully connected layer and a hash layer:

[0085] The third fully connected layer is used to perform operations on the filtered third-branch feature map and output the classification features of the filtered third-branch feature map;

[0086] The hash layer is used to generate a hash code based on the classification features and classification weights of the filtered third-branch feature map, and its expression is as follows:

[0087]

[0088] Among them, represents the hashing operation; avg(·) represents the global average pooling function; W(·) represents converting a high-dimensional feature vector into a low-dimensional one; sign(·) represents the activation function, and w i is the weight, and the purpose of w i is to assign appropriate weights to different elements in the feature vector while filtering out redundant or useless elements, thus helping the network generate a more compact binary code. Its specific settings are as follows:

[0089]

[0090] y i o = fc(avg(f i o ))

[0091] Among them, argmax(·) represents the index of the maximum value in the y i o tensor.

[0092] The hashing network also includes a first fully connected layer and a second fully connected layer;

[0093] The first fully connected layer is used to perform operations on the filtered first branch feature map and output the classification features of the filtered first branch feature map;

[0094] The second fully connected layer is used to perform operations on the filtered second branch feature map and output the classification features of the filtered second branch feature map.

[0095] Preferably, the basic network and the feature filtering module share network parameters. All three branches participate in the network training, but only branch 3 is used during testing. Its main advantage is that fewer parameters are called, thus reducing the inference time.

[0096] Step 120, train the multi-branch multi-loss learning hashing network according to the multi-stage loss function.

[0097] Specifically, using the multi-stage loss function to optimize network training can prompt the generation of a more compact hash code.

[0098] Preferably, the multi-stage loss function includes the significance loss, classification loss, adversarial loss, target loss, and hash loss calculated based on the classification features of the filtered first branch feature map, the classification features of the filtered second branch feature map, and the classification features of the filtered third branch feature map.

[0099] Preferably, the significance loss is calculated based on the classification features of the filtered second-branch feature map to enhance the discriminative ability of the binary code;

[0100] The classification loss is calculated based on the classification features of the filtered first-branch feature map to measure the global semantic information;

[0101] The adversarial loss is calculated based on the classification features of the filtered first-branch feature map and the classification features of the filtered third-branch feature map to optimize the network training process;

[0102] The object loss is calculated based on the filtered third-branch feature map and the classification features of the filtered third-branch feature map to mine discriminative regions;

[0103] The hash loss is calculated based on the hash code to optimize the hash code in binary code learning.

[0104] Specifically, for the significance loss and the classification loss, the widely used cross-entropy loss function is adopted, and its mathematical expression is as follows:

[0105]

[0106] where L sig is the significance loss, L cls is the classification loss, t i represents the true label; represents the j-th element of ce (), and l

[0107] () is the cross-entropy loss function.

[0108] In addition, the Hamming loss is adopted as the adversarial loss to guide the network to learn the expected classification target.

[0109]

[0110] where L obj is the object loss, L mc () is the MC-Loss function.

[0111] The hash loss and its optimization utilize the class proxy-based loss to maintain the consistency of the hash codes between similar samples while reducing the negative impact of intra-class variations.

[0112]

[0113] where L hash is the hash loss, a j ∈R1×k represents the proxy vector of the j-th class; T ∈ R c×m represents the label matrix, and the multi-stage loss function can be expressed as:

[0114] L mtl = L hash + L sig + L cls + β × L adv + γ × L obj

[0115] where L mtl is the multi-stage loss, β and γ represent the weight coefficients, and L adv is the adversarial loss.

[0116] The test network structure only contains Branch 3 and does not involve the other two branches. The main advantage of this design is that the number of its model parameters is moderate. The hash function has the following expression:

[0117]

[0118] Step 130, generate the hash code of the sample image through the trained multi-branch multi-loss learning hash network.

[0119] Step 140, obtain the query image.

[0120] Step 150, generate the hash code of the query image through the trained multi-branch multi-loss learning hash network.

[0121] Step 160, generate the image retrieval result based on the hash code of the sample image and the hash code of the query image.

[0122] Specifically, to generate the image retrieval result, it is necessary to calculate the similarity between the hash code of the query image and the hash code of the sample image. For example, calculate the Hamming distance, cosine similarity, etc. between the hash code of the query image and the hash code of the sample image. According to the calculated similarity metric value, sort the sample images. Select several sample images with the highest similarity metric value as the retrieval result. These sample images are the most similar to the query image in terms of features and are therefore considered to be the images related to the query image.

[0123] To verify the effectiveness of this method, extensive experiments were conducted on 3 fine-grained datasets (Stanford Cars, FGVC-Aircraft, CUB-200-2011) for performance evaluation. The results show that the performance of this method is better than the state-of-the-art general hash method.

[0124] Figure 5It is a schematic diagram of the modules of a fine-grained image retrieval system based on multi-branch multi-loss learning shown in some embodiments of this specification. As Figure 5 shown, the fine-grained image retrieval system based on multi-branch multi-loss learning may include a model establishment module, a model training module, a hash code generation module, an image acquisition module, and an image retrieval module.

[0125] The model establishment module is used to establish a hash network based on multi-branch multi-loss learning. Among them, the hash network based on multi-branch multi-loss learning at least includes a basic network, a feature filtering module, and a hash network. The basic network is used to extract the feature map of the input image. The feature filtering module is used to remove noise features in the channel and spatial dimensions of the feature map and generate a filtered feature map. The hash network is used to generate a hash code based on the filtered feature map;

[0126] The model training module is used to train the hash network based on multi-branch multi-loss learning according to a multi-stage loss function;

[0127] The hash code generation module is used to generate the hash code of the sample image through the trained hash network based on multi-branch multi-loss learning;

[0128] The image acquisition module is used to acquire a query image;

[0129] The hash code generation module is also used to generate the hash code of the query image through the trained hash network based on multi-branch multi-loss learning;

[0130] The image retrieval module is used to generate an image retrieval result based on the hash code of the sample image and the hash code of the query image.

[0131] The fine-grained image retrieval system based on multi-branch multi-loss learning can be used to execute the above-mentioned fine-grained image retrieval method based on multi-branch multi-loss learning, which will not be elaborated here.

[0132] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other deformations may also fall within the scope of this specification. Therefore, as an example rather than a limitation, alternative configurations of the embodiments of this specification can be regarded as consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments clearly introduced and described in this specification.

Claims

1. A fine-grained image retrieval method based on multi-branch and multi-loss learning, characterized in that Including: Construct a multi-branch multi-loss learning hash network, where the multi-branch multi-loss learning hash network at least includes a basic network, a feature filtering module, and a hash network. The basic network is used to extract the feature map of the input image. The feature filtering module is used to remove noise features in the channel and spatial dimensions of the feature map to generate a filtered feature map. The hash network is used to generate a hash code based on the filtered feature map; Train the multi-branch multi-loss learning hash network according to a multi-stage loss function; Generate the hash code of the sample image through the trained multi-branch multi-loss learning hash network; Obtain a query image; Generate the hash code of the query image through the trained multi-branch multi-loss learning hash network; Generate an image retrieval result based on the hash code of the sample image and the hash code of the query image.

2. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 1, characterized in that The basic network includes a feature extraction unit, a dual-selection saliency region erasing unit, and a dual-filtering target localization unit; The feature extraction unit includes a first feature extraction branch, a second feature extraction branch, and a third feature extraction branch. Among them, the first feature extraction branch is used to perform preliminary feature extraction on the image to obtain a preliminary feature map; The dual-selection saliency region erasing unit is used to generate a region-erased image of the input image based on the preliminary feature map; The second feature extraction branch is used to extract the feature map of the region-erased image The dual-filtering target localization unit is used to generate a target localization image of the input image based on the preliminary feature map; The third feature extraction branch is used to extract the feature map of the target localization image. Among them, the feature map of the input image includes the preliminary feature map, the feature map of the region-erased image, and the feature map of the target localization image.

3. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 2, wherein The dual-selection saliency region erasing unit includes a channel filtering group and a spatial erasing group; The channel filtering group is used to calculate the ratio of non-zero pixels in each channel of the preliminary feature map, calculate the average value of the non-zero pixel ratios of all channels, and delete the channels lower than the average value to obtain a channel-filtered feature map; The spatial erasing group is used to perform channel aggregation on the channel-filtered feature map, adjust the preliminary feature map to the same size as the input image through bilinear interpolation, erase the region with the highest response based on a threshold, and multiply it element-wise with the input image to generate a region-erased image of the input image.

4. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 3, wherein The dual-filtering target localization unit includes a spatial filtering group and a noise background filtering group; The spatial filtering group is used to perform channel aggregation on the preliminary feature map, perform an average operation on each channel of the channel-aggregated feature map in the spatial dimension, remove features smaller than the average value, and generate a spatially filtered feature map; The noise background filtering group is used to filter the noise features of the spatially filtered feature map through the largest connected component method, detect the position of the key region, generate a region mask of the same size as the input image based on the position of the key region using bilinear interpolation, and multiply the region mask element-wise with the input image to generate a target localization image of the input image.

5. The fine-grained image retrieval method based on multi-branch multi-loss learning according to any one of claims 2-4, characterized in that, The feature filtering module includes a first feature filtering branch, a second feature filtering branch, and a third feature filtering branch; The first feature filtering branch is used to filter the preliminary feature map to generate the filtered first branch feature map; The second feature filtering branch is used to filter the region erased image of the input image to generate the filtered second branch feature map; The third feature filtering branch is used to filter the target localization image of the input image to generate the filtered third branch feature map, where the filtered feature map includes the filtered first branch feature map, the filtered second branch feature map, and the filtered third branch feature map.

6. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 5, characterized in that The first feature filtering branch includes a channel feature filtering group and a spatial feature filtering group; The channel feature filtering group is used to calculate the proportion of non-zero values in each channel of the preliminary feature map and delete the channels with a proportion less than the proportion threshold to generate the feature map after channel feature filtering; The spatial feature filtering group is used to perform channel aggregation processing on the feature map after channel feature filtering to generate the feature map after aggregation processing, and generate the filtered first branch feature map based on the feature values and the global average value at each spatial position of the feature map after aggregation processing.

7. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 5, wherein The hash network includes at least a third fully connected layer and a hash layer: The third fully connected layer is used to perform operations on the filtered third branch feature map and output the classification features of the filtered third branch feature map; The hash layer is used to generate a hash code based on the classification features and classification weights of the filtered third branch feature map.

8. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 7, wherein The hash network further includes a first fully connected layer and a second fully connected layer; The first fully connected layer is used to perform operations on the filtered first branch feature map and output the classification features of the filtered first branch feature map; The second fully connected layer is used to perform operations on the filtered second branch feature map and output the classification features of the filtered second branch feature map; The multi-stage loss function includes a saliency loss, a classification loss, an adversarial loss, a target loss, and a hash loss calculated based on the classification features of the filtered first branch feature map, the classification features of the filtered second branch feature map, and the classification features of the filtered third branch feature map.

9. The fine-grained image retrieval method based on multi-branch multi-loss learning according to claim 8, characterized in that, The saliency loss is calculated based on the classification features of the filtered second branch feature map; The classification loss is calculated based on the classification features of the filtered first branch feature map; The adversarial loss is calculated based on the classification features of the filtered first branch feature map and the classification features of the filtered third branch feature map; The target loss is calculated based on the filtered third branch feature map and the classification features of the filtered third branch feature map; The hash loss is calculated based on the hash code.

10. A fine-grained image retrieval system based on multi-branch multi-loss learning, characterized in that, A method for fine-grained image retrieval based on multi-branch multi-loss learning for performing any one of claims 1-9 includes: A model building module for building a multi-branch multi-loss learning hashing network, wherein the multi-branch multi-loss learning hashing network at least includes a basic network, a feature filtering module and a hashing network. The basic network is used to extract the feature map of the input image. The feature filtering module is used to remove the noise features in the channel and spatial dimensions of the feature map and generate a filtered feature map. The hashing network is used to generate a hash code based on the filtered feature map; A model training module for training the multi-branch multi-loss learning hashing network according to a multi-stage loss function; A hash code generation module for generating the hash code of the sample image through the trained multi-branch multi-loss learning hashing network; An image acquisition module for acquiring a query image; The hash code generation module is further used to generate the hash code of the query image through the trained multi-branch multi-loss learning hashing network; An image retrieval module for generating an image retrieval result based on the hash code of the sample image and the hash code of the query image.