Fine-grained image retrieval method based on compact representation modeling and semantic tag guidance
By using VisualTransformer feature extraction network and MUSTA module in fine-grained image retrieval, combined with semantic-guided binary encoding module, the problem of difficulty in mining intermediate features of images and learning fine-grained semantics in the prior art is solved, and a more accurate and efficient fine-grained image retrieval is achieved.
Patent Information
- Application Number
- CN202510002116.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing fine-grained image retrieval methods are difficult to efficiently mine the significant and tight intermediate features of the image, it is difficult to construct long-distance dependence between features, and it is difficult to learn fine-grained semantics at the hash task head for discriminating fine-grained categories.
A feature extraction network based on VisualTransformer is adopted to separate global and local representation learning branches, and a MUSTA module is built in the local level feature learning branches to aggregate significant visual word vectors. At the same time, through the semantic-guided fine-grained binary encoding module, the similarity of fine-grained semantic category labels is used to guide the generation of coding centers, and the center loss function is optimized to enhance the model's ability to distinguish fine-grained categories.
It significantly enriches multi-layer local visual information, improves the network's ability to represent images, enhances the ability to discriminate fine-grained categories, and improves the accuracy and efficiency of retrieval.
Smart Images

Figure CN120030182A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fine-grained image analysis, and in particular to a fine-grained image retrieval method based on compact representation modeling and semantic label guidance. Background Art
[0002] The fine-grained image retrieval task aims to accurately retrieve images that belong to fine-grained subcategories rather than traditional general categories. This task has a wide range of application scenarios in real life, such as smart surveillance, smart retail, and smart transportation, and has therefore received increasing attention in the field of artificial intelligence technology.
[0003] Existing traditional fine-grained image retrieval work can be divided into fine-grained content-based image retrieval (FG-CBIR) and fine-grained sketch-based image retrieval (FG-SBIR) according to the modality of the query image. In terms of FG-CBIR, the Selective Convolutional Descriptor Aggregation (SCDA) method first attempted to use a pre-trained convolutional neural network (CNN) to select meaningful deep descriptors by locating salient objects without ordering image labels or bounding box annotations. Subsequently, supervised methods were developed using metric learning and additional submodules customized for fine-grained objects, effectively breaking through the accuracy limitations of unsupervised retrieval. The Decorrelated Global-aware Centralized Ranking Loss (DGCRL) method eliminates the gap between the inner product and the Euclidean distance by adding a normalized scale layer in the training and testing stages to enhance intra-class separability and inter-class compactness. While FG-SBIR aims to use the retracted sketch as a query to match a specific image instance, existing methods usually design and train a joint embedding space to bridge the gap between the sketch-image domain. The above-mentioned fine-grained image retrieval methods all focus on establishing discriminative fine-grained representations, but ignore the retrieval time and feature storage consumption, which greatly hinders the retrieval application on large databases.
[0004] Learning-based fine-grained hash coding can effectively solve the above problems. It maps the high-dimensional real-valued features of the image into low-dimensional binary codes, so that the search results in the coding space are as close as possible to the results in the original space, thereby significantly reducing the storage cost and speeding up the retrieval process. The general hashing algorithm focuses the network's attention on the global features of the object. However, with the emergence of large-scale fine-grained datasets, most general hashing methods are difficult to meet the needs of fine-grained image analysis tasks due to the lack of mining of fine-grained features of the object. Therefore, many hashing methods for fine-grained samples have been kicked out to effectively solve the high computational cost while extracting fine-grained features in large-scale fine-grained datasets. The ExchNet method adds additional modules to extract local features to represent the parts of the visual object, and fuses global and local features to generate a unified binary code. The Suppression-Enhancing Mask based Attention and Interactive Channel Transformation (SEMICON) method proposes a suppression-enhancing mask based attention operation to maintain the relationship between different activation regions, and combines it with a two-step interactive channel transformation module to calculate the attention between different channels to effectively capture the correlation between fine-grained parts, and finally generate binary encoding for the task. The Attributes Grouping and Mining Hashing (AGMH) method groups and embeds category-specific visual attributes into multiple descriptors, optimizes them through attention-distributed loss to generate comprehensive feature representations, and proposes step-by-step interactive external attention to mine key attributes in each descriptor, while building correlations between fine-grained attributes and objects.
[0005] However, the above methods are based on CNN using the local attention mechanism to extract subtle discriminative visual features of images. It is difficult to efficiently mine significant and compact intermediate features of images, difficult to build long-distance dependencies between features in the backbone network, and difficult to learn fine-grained semantics in the hash task head for discriminating fine-grained categories. Therefore, there is a need in the field of technology for an improved image retrieval method that overcomes the above defects. Summary of the invention
[0006] In view of the technical defects existing in the above-mentioned fine-grained image retrieval method, the object of the present invention is to propose a fine-grained image retrieval method based on compact representation modeling and semantic label guidance. The method is built on the ViT (Visual Transformer) backbone network, and two representation learning branches are subsequently separated, namely a global branch for learning target-level features and a local branch for learning part-level features, and further integrated to generate an overall compact representation. However, the traditional ViT is designed for general visual tasks and is difficult to explore and integrate more subtle visual areas. In order to overcome this limitation, a MUSTA module (Multi-Level SalientToken Aggregating, multi-level significant visual word vector aggregation module) is constructed in the local level feature learning branch, and the most significant visual word vectors are selected according to the attention weights of the low-level and high-level visual word vectors in the backbone network, and further aggregated to the last Transformer layer, thereby significantly enriching the multi-layer local visual information. In addition, a semantic-guided fine-grained binary coding (SFBC) module is established to guide the generation of encoding centers by considering the similarity between fine-grained semantic category labels. At the same time, the center loss function is weighted by similarity, so that the model can pay more attention to the significant differences between samples belonging to similar categories.
[0007] According to an embodiment of the present invention, a fine-grained image retrieval method based on compact representation modeling and semantic label guidance is provided, comprising the following steps:
[0008] Step S1, preparing an image dataset containing fine-grained categories, performing data preprocessing and data division, and obtaining a preprocessed image dataset;
[0009] Step S2, constructing a feature extraction network, including constructing a feature extraction network based on a multi-layer visual Transformer as a backbone network, for extracting visual features of images in the preprocessed image data set to obtain visual word vectors, the feature extraction network also includes a global level feature learning branch and a local level feature learning branch, wherein the local level feature learning branch includes a MUSTA module;
[0010] Step S3, generating a hash coding center, building a bag-of-words model based on the fine-grained image labels, and counting the frequency of each word appearing in the category label as the word frequency, calculating the semantic similarity matrix between all category labels based on the statistically obtained word frequency, and finally generating the coding center of each category with the semantic similarity matrix as the weight;
[0011] Step S4, constructing a compact coding mapping unit, inputting the visual word vector extracted by the feature extraction network into the constructed coding mapping unit, and using the activation function in the coding mapping unit to map the high-dimensional real-valued features of the visual word vector into low-dimensional binary codes, to obtain a compact representation of the image, and constructing a retrieval coding library with the obtained compact representation of the image;
[0012] Step S5, training the constructed feature extraction network and compact coding mapping unit to obtain trained feature extraction network and compact coding mapping unit;
[0013] Step S6, perform fine-grained image retrieval, input the image to be detected into the trained feature extraction network and the compact coding mapping unit to obtain the high-dimensional features of the image to be detected, map the high-dimensional features of the image to be detected into low-dimensional codes as the codes to be detected, match the obtained codes to be detected with the codes in the retrieval code library, and the sample with the smallest distance is the final retrieval result.
[0014] Optionally, step S2 specifically includes the following steps:
[0015] Step S2.1, constructing a feature extraction network, first constructing a multi-layer visual Transformer including a global level feature learning branch and a local level feature learning branch as the backbone network of the feature extraction network, wherein the local level feature learning branch includes a MUSTA module;
[0016] Step S2.2, input the preprocessed image dataset and use the constructed multi-layer visual TransformerΦ ViT (·) Extract a sequence of global-level visual word vectors from image I in the input preprocessed image dataset:
[0017] T gl =Φ ViT (I);
[0018] Step S2.3, in parallel, the set MUSTA module is used to select several relatively significant visual word vectors from each attention head of each Transformer layer in all Transformer layers except the last layer of the multi-layer visual Transformer, and all the significant visual word vectors selected from each Transformer layer except the last layer are aggregated into the final input sequence Z′ L-1 , where L represents the total number of Transformer layers included in the multi-layer visual Transformer;
[0019] Step S2.4: The final input sequence Z′ after aggregation L-1Input into the last Transformer layer in the local-level feature learning branch, and through processing, a local-level visual word vector sequence is obtained:
[0020]
[0021] Optionally, step S2.3 specifically includes the following steps:
[0022] Step S2.3.1, the total number of Transformer layers included in the multi-layer visual Transformer is L layers, and the input of the last Transformer layer in the local-level feature learning branch is Z′ L-1 , and the attention weights of each layer of the first L - 1 layers of Transformer layers, each with M attention heads, are expressed as:
[0023]
[0024] Among them, A l represents the attention weight of the l-th Transformer layer, l represents the current Transformer layer number and 1 ≤ l ≤ L - 1, represents the attention weight of the m-th attention head in A l , represents the attention weight corresponding to the N T -th visual word vector in, m represents the current attention head number in this Transformer layer, N T represents the number of non-class visual word vectors of each attention head;
[0025] Step S2.3.2, in the MUSTA module, multiply the attention weights layer by layer from the L-th Transformer layer and before to obtain the fused attention weight used to record the forward propagation of visual word vectors from the input layer to higher layers:
[0026]
[0027] Among them, a′ represents the fused attention weight, w represents the number of the Transformer layer, and A w represents the attention weight of the w-th Transformer layer and 1 ≤ w ≤ l;
[0028] Step S2.3.3, select the attention weights corresponding to the class visual word vectors from the fused attention weight a′:
[0029]
[0030] And screen out the indices corresponding to the n′ visual word vectors with the highest weight values from each attention head:
[0031]
[0032] Among them, v l represents the fused attention weight corresponding to the category visual word vector in the lth layer Transformer layer, Indicates v l The attention weight of the mth attention head in , K l Represents the index values corresponding to several visual word vectors with the highest fused attention weights in the lth layer Transformer layer, K l The visual word vector index value of the mth attention head in . The TopSelection(·) function means using the max(·) function to filter the index corresponding to the highest attention weight.
[0033] Step S2.3.4, in the output sequence of each Transformer layer of the first L-1 layers of Transformer layers, find the index value K that matches the selected index value of that layer l The corresponding visual word vector:
[0034] T s,l =T l [K l ]
[0035] Among them, T l Represents the output sequence of the lth layer Transformer layer;
[0036] Step S2.3.4, aggregate the visual word vectors of the output sequences of each Transformer layer of the first L-1 layers of Transformer layers obtained above with the category visual word vectors of the original input sequence of the Lth Transformer layer through the MUSTA module as the final input sequence of the last Transformer layer in the local level feature learning branch:
[0037]
[0038] in, Represents the category visual word vector of the original input sequence of the L-th Transformer layer.
[0039] Optionally, the step S3 specifically includes the following steps:
[0040] Step S3.1, generate a set of q binary centers with at least distance d between K subcategories:
[0041]
[0042] Where k represents the fine-grained category number, c k represents the hash code center of the kth class, d is the minimum distance between any two representation centers obtained based on the Gilbert-Varshamov boundary theory, K represents the total number of fine-grained categories, and q represents the hash code length;
[0043] Step S3.2, construct a category semantic similarity matrix, adopt a bag-of-words model containing w' words, and represent the K fine-grained labels as a feature matrix V based on word occurrence probability. The similarity matrix is expressed as:
[0044] S=V·V T ;
[0045] Based on the category similarity matrix S, the above minimum distance d is replaced by a fine-grained weighted distance constraint d between two categories. i,j :
[0046] d i,j =(1+S i,j )·d;
[0047] Then, the hash code centers of each fine-grained category are generated by solving the following optimization problem:
[0048]
[0049] st||c i -c j || H ≥d i,j ,1≤i,j≤K,i≠j,
[0050] in,‖·‖ H is the Hamming distance, i and j represent the fine-grained category numbers, and c K 、c i 、c j Respectively represent the hash code center of the Kth, ith, and jth categories, d i,j represents the fine-grained weighted distance constraint between the i-th class and the j-th class, S i,j It represents the similarity between the i-th and j-th labels.
[0051] Optionally, the step S4 specifically includes the following steps:
[0052] Step S4.1, by respectively embedding the visual word vector sequence T at the global level gl and the local level visual word vector sequence T loc Apply global average pooling to obtain the global level real-valued feature x gl and local level real-valued features x loc ;
[0053] Step S4.2, use the global level linear mapping matrix W gl and the local level linear mapping matrix W loc , respectively, the high-dimensional global level real-valued features x gl and high-dimensional local-level real-valued features x loc Mapped to low-dimensional global level real-valued features h gl and low-dimensional local level real-valued features h loc ;
[0054] Step S4.3, transform the low-dimensional global level real-valued feature h gl and low-dimensional local level real-valued features h loc Mapping to low-dimensional compact encoding b = {b gl ,b loc}, we get a compact representation of the image, where b gl represents the global level low-dimensional compact encoding, b loc Represents a local-level low-dimensional compact encoding.
[0055] Step S4.4, the low-dimensional compact codes corresponding to the samples of the images in all the input preprocessed image data sets are stored according to the image sample numbers in each gallery, and a retrieval code library with an index set of Ω is constructed.
[0056] Optionally, step S5 specifically includes the following steps:
[0057] Step S5.1, embed the similarity matrix S into the compact code b of the nth sample of the kth class n and the corresponding k-th hash code center c k , we get the similarity between the compact code and its corresponding hash code center:
[0058]
[0059] Among them, Sim(·,·) represents the scaled cosine similarity calculation, P n,k Represents the compact code b n and the corresponding k-th hash code center c k similarity, t represents the fine-grained category number, K represents the total number of fine-grained categories, S k,t Indicates the similarity between the k-th and t-th class labels, c t represents the hash code center of the tth class;
[0060] Step S5.2, establish the semantic similarity weighted center loss function:
[0061]
[0062] Among them, Z n,kis an indicator factor. If the nth input image I n The compact encoding of n belongs to the kth category, then Z n,k =1, otherwise Z n,k =0;
[0063] Step S5.3, establish quantization loss and hash coding loss function:
[0064]
[0065] Among them, λ 1 , 2 represents the loss term coefficient, u represents the image sample number in the gallery, v represents the query sample number, Ω represents the index set of gallery samples in the training data, Γ represents the index set of query samples in the training data, and b u , b v Represents the low-dimensional hash code of samples u and v, respectively, u represents the low-dimensional real-valued encoding of sample u, L uv represents paired sample supervision information and L uv ∈{-1,+1} (Q×D) , Q represents the number of query samples, D represents the number of gallery samples, The first term is to convert the real-valued low-dimensional feature vector h u The second term is the general semantic preservation loss.
[0066] Step S5.4, the quantization loss and hash coding loss function Similarity with semantic weighted center loss function Combined, we finally get the total loss function that optimizes the entire network:
[0067]
[0068] Step S5.5: The final total loss function includes the losses of the feature extraction network and the compact coding mapping unit. During the gradient back-propagation process, the network parameters of the feature extraction network and the compact coding mapping unit are updated at the same time until the loss of the loss function converges, thus obtaining a trained feature extraction network and a compact coding mapping unit.
[0069] Compared with the prior art, the fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention has at least the following beneficial effects.
[0070] 1. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided by the present invention makes up for the defect of insufficient modeling of significant visual information in existing algorithms and expands the idea of fine-grained image retrieval.
[0071] 2. The salient visual features of the image are extracted through the Multi-Level Salient Token Aggregating (MUSTA) module, which enables the backbone network to capture rich fine-grained visual information and improves the network's ability to represent images.
[0072] 3. Through the semantic-guided fine-grained binary coding (SFBC) module, the similarity information between fine-grained category labels is integrated into the encoding center generation and center loss function, so that the encoding mapping layer can distance the encodings of samples belonging to similar fine-grained categories, reducing the difficulty of distinguishing difficult samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. The features and advantages of the present invention can be more clearly understood by referring to the drawings. The drawings are schematic and should not be understood as limiting the present invention in any way. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0074] Figure 1 The present invention provides a flowchart of a fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to an embodiment of the present invention.
[0075] Figure 2 A schematic diagram of the overall structure of a model constructed by a fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention.
[0076] Figure 3 A schematic diagram of the MUSTA module structure in a fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention.
[0077] Figure 4 A schematic diagram of the structure of the SFBC module in the fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention.
[0078] Figure 5 The figure is a schematic diagram of the specific process of the MUSTA module in the fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0079] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0080] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0081] The fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided according to an embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0082] Explanation of symbols:
[0083] I represents the image in the input preprocessed image dataset;
[0084] T gl Represents a sequence of global-level visual word vectors;
[0085] L represents the total number of Transformer layers included in the multi-layer visual Transformer;
[0086] Z′ L-1 represents the final input sequence obtained by aggregating all the salient visual word vectors selected in each Transformer layer except the last layer;
[0087] M represents the number of attention heads;
[0088] A l represents the attention weight of the lth Tranformer layer;
[0089] l represents the current Transformer layer number, 1≤l≤L-1;
[0090] Indicates A l The attention weight of the mth attention head in ;
[0091] express N T The attention weights corresponding to the visual word vectors;
[0092] m represents the current attention head number in the current Tranformer layer;
[0093] N T Represents the number of non-category visual word vectors for each attention head;
[0094] a′ represents the fusion attention weight;
[0095] A w represents the attention weight of the w-th Tranformer layer;
[0096] w represents the number of the Transformer layer, 1≤w≤l;
[0097] v l represents the fused attention weight corresponding to the category visual word vector in the lth layer Transformer layer;
[0098] Indicates v l The attention weight of the mth attention head in ;
[0099] K l Represents the index values corresponding to several visual word vectors with the highest fused attention weights in the lth Transformer layer;
[0100] K l The visual word vector index value of the mth attention head in ;
[0101] The TopSelection(·) function means using the max(·) function to filter the index corresponding to the highest attention weight;
[0102] n' represents the number of visual word vectors with the highest weight value in each attention head;
[0103] T l Represents the output sequence of the lth layer Transformer layer;
[0104] Represents the category visual word vector of the original input sequence of the L-th layer Transformer layer;
[0105] T loc Represents a sequence of local-level visual word vectors;
[0106] Z L-1 represents the original complete input sequence;
[0107] c k represents the hash code center of the kth class;
[0108] k represents the fine-grained category number;
[0109] ‖·‖ H is the Hamming distance;
[0110] K represents the total number of fine-grained categories;
[0111] i and j represent the fine-grained category numbers;
[0112] c K 、c i 、c j Represent the hash code centers of the Kth, ith, and jth categories respectively;
[0113] q represents the hash code length;
[0114] d i,j represents the fine-grained weighted distance constraint between the i-th class and the j-th class;
[0115] S i,j Indicates the similarity between the i-th and j-th class labels;
[0116] x gl Represents high-dimensional global-level real-valued features;
[0117] x loc Represents high-dimensional local-level real-valued features;
[0118] W gl Represents the global level linear mapping matrix;
[0119] W loc Represents the local level linear mapping matrix;
[0120] h gl Represents low-dimensional global-level real-valued features;
[0121] h loc Represents low-dimensional local-level real-valued features;
[0122] b represents low-dimensional compact coding;
[0123] b gl Represents a low-dimensional compact encoding at the global level;
[0124] b loc Represents low-dimensional compact encoding at the local level;
[0125] S represents the similarity matrix;
[0126] Sim(·,·) represents the scaled cosine similarity calculation;
[0127] P n,k Represents the compact code b n and its corresponding hash code center c k similarity;
[0128] t represents the fine-grained category number;
[0129] S k,t Indicates the similarity between the k-th and t-th class labels;
[0130] c t represents the hash code center of the tth class;
[0131] Z n,k is the indicator factor;
[0132] I n represents the nth input image;
[0133] b n represents the compact encoding of the nth sample of the kth class;
[0134] c k′ Indicates the encoding center of other categories with similar semantics to the kth category;
[0135] represents the semantic similarity weighted center loss function;
[0136] represents the hash coding loss function;
[0137] represents the total loss function;
[0138] λ 1 , 2 represents the loss term coefficient;
[0139] u represents the image sample number in the gallery;
[0140] v represents the query sample number;
[0141] Ω represents the index set of gallery samples in the training data;
[0142] Γ represents the index set of query samples in the training data;
[0143] b u , b v Represent the low-dimensional hash codes of samples u and v respectively;
[0144] h u Represents the low-dimensional real-valued encoding of sample u;
[0145] L uv Represents sample supervision information;
[0146] Q represents the number of query samples;
[0147] D represents the number of gallery samples.
[0148] like Figure 1 As shown, the fine-grained image retrieval method based on compact representation modeling and semantic label guidance provided in accordance with an embodiment of the present invention includes the following steps.
[0149] Step S1, prepare an image dataset containing fine-grained categories, perform data preprocessing and data partitioning, and obtain a preprocessed image dataset. Optionally, the image dataset containing fine-grained categories in this step can be selected from a public database, and the data preprocessing method includes random cropping, random flipping, random scaling, etc. The data partitioning follows the data partitioning rules of each public dataset.
[0150] Step S2: construct a feature extraction network. Figure 2 As shown, a feature extraction network is constructed, including a multi-layer visual Transformer based on ViT (Visual Transformer) as the backbone network, and visual features are extracted from the images in the preprocessed image dataset to obtain visual word vectors. Visual features include global level features (i.e., target features) and local level features (i.e., component features). The multi-layer visual Transformer of the feature extraction network includes a global level feature learning branch and a local level feature learning branch, which model the global level features and the local level features respectively. In addition, the local level feature learning branch also includes a MUSTA module (multi-level salient visual word vector aggregation module), as shown in FIG. Figure 3 As shown in FIG. 1 , the MUSTA module is used to filter the significant visual word vectors of several Transformer layers of the visual Transformer, and fuse the filtered significant visual word vectors as the input of the last Transformer layer to enhance the saliency of features in the local level feature learning branch.
[0151] Furthermore, the feature extraction network constructed in step S2 is a visual Transformer including a MUSTA module. Specifically, for fine-grained tasks, both global-level features and local-level features are very important. Figure 2 , Figure 3 and Figure 5 As shown, step S2 specifically includes the following steps.
[0152] Step S2.1, construct a feature extraction network. First, construct a multi-layer visual Transformer including a global level feature learning branch and a local level feature learning branch as the backbone network of the feature extraction network, wherein the local level feature learning branch includes a MUSTA module.
[0153] Step S2.2, input the preprocessed image dataset and use the constructed multi-layer visual Transformer, φ ViT (·) to extract the global level visual word vector sequence of image I in the input preprocessed image dataset: T gl =Φ ViT (I).
[0154] Step S2.3, at the same time, the set MUSTA module is used in parallel to select several relatively significant visual word vectors from each attention head of each Transformer layer in all Transformer layers except the last Transformer layer of the multi-layer visual Transformer, and all the significant visual word vectors selected from each Transformer layer except the last layer are aggregated into the final input sequence Z′ L-1 , to enrich the discriminative visual cues. Where L represents the total number of Transformer layers included in the multi-layer visual Transformer.
[0155] Furthermore, the MUSTA module of the local level feature learning branch in step S2 extracts and aggregates the most significant fine-grained features (i.e., more significant visual word vectors) in all low layers and high layers except the last Transformer layer in the multi-layer visual Transformer network, so as to enhance the significance of the features learned by the local level feature learning branch and reduce the redundant information in the visual word vector sequence.
[0156] Step S2.3 specifically includes the following steps.
[0157] Step S2.3.1: In order to make full use of the attention information, more significant information is input into the last Transformer layer. In this step, the attention weights of each Transformer layer are fused and assigned to each visual word vector in the sequence to establish the association between the attention weight and the importance of the word vector. Set the input of the last Transformer layer in the local level feature learning branch to Z′ L-1 , then the attention weights of each layer of the first L-1 layers of Tranformer with M attention heads are expressed as:
[0158]
[0159] Among them, M represents the number of attention heads, A 1 represents the attention weight of the lth Tranformer layer, Indicates A l The attention weight of the mth attention head in , express N T The attention weight corresponding to the visual word vector, l represents the current Transformer layer number and 1≤l≤L-1, m represents the current attention head number in the Transformer layer, N T Represents the number of non-class visual word vectors for each attention head.
[0160] Step S2.3.2, in the MUSTA module, the attention weights of the lth Tranformer layer and before are multiplied layer by layer to obtain the fused attention weights used to record how the visual word vectors are forward propagated from the input layer to higher layers:
[0161]
[0162] Among them, a′ represents the fusion attention weight, w represents the number of the Transformer layer, and A w represents the attention weight of the w-th Tranformer layer and 1≤w≤l.
[0163] In step S2.3.3, the original single-layer attention weight a is more suitable for selecting discriminative regions. Since the category visual word vector can represent the global level features of the image, the weight corresponding to the category visual word vector is selected from the fused attention weight a′:
[0164]
[0165] And filter out the indexes corresponding to the n' visual word vectors with the highest weight values from each attention head:
[0166]
[0167] Among them, v l represents the fused attention weight corresponding to the category visual word vector in the lth layer Transformer layer, Indicates v l The attention weight of the mth attention head in , K l Represents the index values corresponding to several visual word vectors with the highest fused attention weights in the lth layer Transformer layer, K l The visual word vector index value of the mth attention head in , and the TopSelection(·) function means using the max(·) function to filter the indexes corresponding to the highest n′ attention weights.
[0168] Step S2.3.4, then, Figure 2 and Figure 5 As shown, in the output sequence of each Transformer layer of the first L-1 layers, find the visual word vector corresponding to the index value corresponding to the visual word vector with the highest fusion attention weight value of the layer:
[0169] T s,l =T l [K l ]
[0170] Among them, T l Represents the output sequence of the lth Transformer layer.
[0171] Step S2.3.4, through the MUSTA module, aggregate the visual word vectors of the output sequences of each Transformer layer of the first L-1 layers of Transformer layers obtained in the above steps with the category visual word vectors of the original input sequence of the Lth Transformer layer as the final input sequence of the last Transformer layer in the local level feature learning branch:
[0172]
[0173] in, Represents the category visual word vector of the original input sequence of the L-th Transformer layer.
[0174] By taking the original complete input sequence Z L-1l Replace it with the most significant visual word vector sequence Z′ obtained above as the final input sequence L-1 , and the category word vector of the original input sequence of the Lth layer Transformer layer Connected as input to the last Transformer layer, it can not only preserve the global level information, but also capture the subtle differences between different subcategories, while removing redundant information such as common features belonging to the same parent category.
[0175] Step S2.4, then, the final input sequence Z′ after aggregation L-1 Input into the last Transformer layer in the local level feature learning branch, and processed to obtain the local level visual word vector sequence
[0176] Step S3, generate a hash code center. Figure 1As shown, the SFBC module (Semantic-Guided Fine-Grained Binary Coding) first builds a bag-of-words model based on the image fine-grained labels to count the frequency of each word in the category label. Subsequently, the semantic similarity matrix between all category labels is calculated based on the word frequency, and finally the encoding center of each category is generated with the semantic similarity matrix as the weight, so that the distance between similar category encoding centers is larger. The sum of the generated encoding centers of each category is the hash encoding center set. Among them, the image fine-grained labels are all category labels contained in the fine-grained dataset, the words are all words in all category labels contained in the fine-grained dataset, and the category labels are all category labels contained in the fine-grained dataset. They are all information that comes with the dataset, and the labels can be loaded directly from each dataset folder.
[0177] Furthermore, the hash code center in step S3 is a global similarity measure, which is used to force the hash codes of similar images to be close to their common hash code center and keep a distance from other hash code centers. Step S3 specifically includes the following steps.
[0178] Step S3.1, first generate a set of q binary centers with at least distance d between K subcategories:
[0179]
[0180] Where k represents the fine-grained category number, c k represents the hash code center of the kth class. And d is the minimum distance between any two representation centers obtained based on the Gilbert-Varshamov boundary theory.
[0181] Step S3.2, then, by taking the compact encoding b of the nth image from the kth class n Zoom in on the corresponding class center c k And at the same time push away the center of other classes, thereby achieving the learning of compact encoding.
[0182] In addition, in order to more accurately distinguish difficult samples, that is, samples from different fine-grained categories of the same parent category with a large number of similar features, a semantically guided fine-grained compact representation learning module can be constructed to impose stricter distance constraints and weights on semantically similar categories than those dissimilar categories to enhance the learning of discriminative compact representations for fine-grained retrieval. The sample refers to the input image to be processed, that is, the image in the preprocessed image dataset.
[0183] The step S3.2 may also include the following steps. First, in order to obtain the semantic similarity between classes, a class semantic similarity matrix is constructed, a bag-of-words model containing w′ words is adopted, and K fine-grained labels are represented as a feature matrix V based on word occurrence probability. Therefore, the similarity matrix is represented as
[0184] S=V·V T .
[0185] Based on the category similarity matrix S, the original distance constraint d is replaced by a fine-grained weighted distance constraint
[0186] d i,j =(1+S i,j )·d
[0187] Then, the hash code centers of each fine-grained category are generated by solving the following optimization problem:
[0188]
[0189] st||c i -c j || H ≥d i,j ,1≤i,j≤K,i≠j,
[0190] in,‖·‖ H is the Hamming distance, K represents the total number of fine-grained categories, i and j represent the fine-grained category numbers, and c K 、c i 、c j Respectively represent the hash code center of the Kth, ith, and jth categories, q represents the hash code length, and d i,j represents the fine-grained weighted distance constraint between the i-th class and the j-th class, S i,j Represents the similarity between the i-th and j-th class labels. Through the above calculation formula, semantically similar categories are given larger distance constraints, thus forcing the corresponding encoding centers to be further away.
[0191] Step S4, construct a compact coding mapping unit. Figure 1 and Figure 4 As shown, all images in the preprocessed image dataset are respectively input into the global level feature learning branch and the local level feature learning branch of the feature extraction network to extract visual word vectors, and the extracted visual word vectors are input into the encoding mapping unit, and an activation function is used to map the high-dimensional real-valued features of the visual word vectors into low-dimensional binary codes to obtain a compact representation of the image, and a retrieval encoding library is constructed using the compact representation of the image.
[0192] Furthermore, the global level visual word vector sequence and the local level visual word vector sequence extracted by step S2 for all images in the retrieval library are mapped into low-dimensional binary codes by the compact coding mapping unit in step S4, and a retrieval coding library is constructed. Figure 4 , specifically, step S4 includes the following steps.
[0193] Step S4.1, by respectively embedding the visual word vector sequence T at the global level gl and the local level visual word vector sequence T loc Applying global average pooling on the , we can obtain high-dimensional global level real-valued features x gl and high-dimensional local-level real-valued features x loc .
[0194] Step S4.2, then use the global level linear mapping matrix W gl and the local level linear mapping matrix W loc , respectively, the high-dimensional global level real-valued features x gl and high-dimensional local-level real-valued features x loc Mapped to low-dimensional global level real-valued features h gl and low-dimensional local level real-valued features h loc .
[0195] Step S4.3, transform the low-dimensional global level real-valued feature h gl and low-dimensional local level real-valued features h loc Mapping to low-dimensional compact encoding b = {b gl ,b loc}. Among them, b gl represents the global level low-dimensional compact encoding, b loc In step S4.3, the low-dimensional global / local level real-valued features are mapped to the global / local level low-dimensional compact codes by an activation function. Optionally, the activation function is a Tanh function.
[0196] Step S4.4, the low-dimensional compact codes corresponding to the samples of the images in all the input preprocessed image data sets are stored according to the image sample numbers in the library, and a retrieval code library with an index set of Ω is constructed.
[0197] Step S5, training the constructed feature extraction network and compact coding mapping unit to obtain trained feature extraction network and compact coding mapping unit. Setting a loss function, using a back propagation algorithm, iteratively updating and optimizing network parameters until the loss of the loss function converges.
[0198] Furthermore, the loss function in step S5 includes a semantic similarity weighted center loss function, a quantization loss, and a hash coding loss. The semantic similarity weighted center loss function is used to shorten the distance to the subordinate subcategory coding center while increasing the distance to the other subcategory coding center, the quantization loss is used to reduce the difference between the high-dimensional real-valued features and the low-dimensional compact representation, and the hash coding loss is used to shorten the intra-class coding distance while increasing the inter-class coding distance. Step S5 specifically includes the following steps.
[0199] Step S5.1, specifically, in order to further preserve the fine-grained semantic information encoded by the learned mechanism, the similarity matrix S is embedded into the compact encoding b of the nth sample of the kth class n and the corresponding k-th hash code center c k , we get the similarity between the compact code and its corresponding hash code center:
[0200]
[0201] Among them, Sim(·,·) represents the scaled cosine similarity calculation, P n,k Represents the compact code b n and the similarity of the hash code center of its corresponding Kth category, t represents the fine-grained category number, K represents the total number of fine-grained categories, S k,t Indicates the similarity between the k-th and t-th class labels, c t Represents the hash code center of the t-th class.
[0202] Step S5.2, establish the semantic similarity weighted center loss function expressed as:
[0203]
[0204] Among them, Z n,k is an indicator factor. If the nth input image I n The compact encoding of n belongs to the kth category, then Z n,k =1, otherwise Z n,k = 0. When minimizing the center loss, the compact encoding b n and the encoding centers c of other categories with similar semantics to category k k′ are given higher penalty weights, thus forcing them to be pushed far away, which greatly helps to learn more discriminative compact encodings to separate input samples from fine-grained categories.
[0205] Step S5.3, establish quantization loss and hash coding loss function:
[0206]
[0207] Among them, λ1 , 2 represents the loss term coefficient, u represents the image sample number in the gallery, Ω represents the index set of gallery samples in the training data, v represents the query sample number, Γ represents the index set of query samples in the training data, and b u , b v Represents the low-dimensional hash code of samples u and v, respectively, u represents the low-dimensional real-valued encoding of sample u, L uv represents sample supervision information, L uv ∈{-1,+1} (Q×D) represents paired supervision information, Q represents the number of query samples, and D represents the number of gallery samples. The first term is to convert the real-valued low-dimensional feature vector h u The second term is the general semantic preservation loss. 1 and λ 2 Can be set to 1 and 200 respectively.
[0208] Step S5.4, the quantization loss and hash coding loss function and semantic similarity weighted center loss function (i.e., fine-grained semantic guided loss) Combined, we finally get the total loss function that optimizes the entire network:
[0209]
[0210] Step S5.5: The final total loss function includes the loss of the feature extraction network and the compact coding mapping unit. During the gradient back propagation process, the network parameters of the feature extraction network and the compact coding mapping unit are updated at the same time until the loss of the loss function converges to obtain the trained feature extraction network and the compact coding mapping unit. The updated network parameters may include the Transformer layer structure in the feature extraction network, the MUSTA module, and the fully connected layer of the compact coding mapping unit.
[0211] Step S6, performing fine-grained image retrieval, inputting the image to be detected into the trained feature extraction network and the compact coding mapping unit to obtain the high-dimensional features of the image to be detected, mapping the high-dimensional features of the image to be detected into low-dimensional codes as the codes to be detected, matching the obtained codes to be detected with the codes in the retrieval code library, and the sample with the smallest distance is the final retrieval result. The method of the present invention can be applied to search engines, e-commerce, social media, etc., and the final retrieval result obtained is the sample with the highest similarity to the image to be detected (i.e., the sample with the smallest hash code distance), which is used as the retrieval result in the application, thereby providing an efficient and accurate image retrieval function.
[0212] Example 1
[0213] A specific embodiment 1 of the fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to the present invention is as follows, including the following steps.
[0214] S1, Dataset preparation. Perform dataset selection, data preprocessing and data partitioning.
[0215] S1.1, this example selects the large-scale fine-grained image CUB200-2011 dataset, Aircraft dataset, Food101 dataset, NABirds dataset and VegFru dataset as fine-grained image retrieval datasets for verifying the invention.
[0216] S1.2, data preprocessing includes image enhancement and image normalization. Specifically, image enhancement includes resizing, random inversion and random cropping, etc. This example selects two enhancement methods: resizing the image to 224×224 pixels and 50% probability random horizontal inversion.
[0217] S1.3, data division is based on the criteria given in each dataset. The CUB200-2011 dataset contains 11,788 images from 200 bird species, 5,994 images for training, and 5,794 images for testing; the Aircraft dataset contains 10,000 images from 100 aircraft species, 6,667 images for training, and 3,333 images for testing; the Food101 dataset contains 101,000 images from 101 foods, 750 images per category for training, and 250 images for testing; the NABirds dataset contains 48,562 images of North American birds from 555 categories, 23,929 images for training, and 24,633 images for testing. The VegFru dataset contains 200 vegetables and 92 fruits, 29,200 images for training, and 116,931 images for testing. For the first three datasets, 2000 images are randomly selected for training in each iteration, while for the last two datasets, 4000 images are randomly selected for training in each iteration.
[0218] S2, builds a feature extraction network to extract global and local features of images in the training set.
[0219] S2.1, ViT-B / 16 is used as the feature extraction network, where the number of network layers, the number of attention heads, and the dimension of the hidden layer feature vector are 12, 12, and 768 respectively. After preprocessing the input training data, the image is divided into several 16×16 pixel non-overlapping image blocks and embedded into a sequence, which is then concatenated with the classified visual word vectors and fused with the positional encoding embedding to form the input visual word vector module as the backbone network input of ViT (Visual Transformer).
[0220] S2.2, for the global level feature learning branch of the backbone network, the visual word vector is directly input into the original ViT (Visual Transformer) network, and the global visual word vector sequence is obtained after 12 Transformer layers.
[0221] S2.3, for the local level feature learning branch of the backbone network, the MUSTA module is introduced into the original ViT network structure to select the three most significant visual word vectors according to the attention weights of the visual word vectors in each Transformer layer and aggregate them as the input sequence of the last layer. The specific model design refers to the specific description in the previous article and will not be repeated here.
[0222] S3, generate hash coding centers. First, according to the fine-grained category labels of the data set, a bag-of-words model that counts the frequency of word occurrences is constructed. Then, a fine-grained semantic similarity matrix is calculated based on the word frequency. Finally, the coding center generation process is weighted to obtain a semantically guided coding center matrix, where the number of rows corresponds to the number of categories in the data set, and the number of columns corresponds to the final coding length of the current instance, which can be set to 12 bits, 24 bits, 32 bits, and 48 bits.
[0223] S4, construct a compact coding mapping unit. The visual word vectors extracted by the global level feature learning branch and the local level feature learning branch of all images in the training dataset are input into the coding mapping unit, and the high-dimensional real-valued features are mapped to low-dimensional binary codes using an activation function to obtain a compact representation of the image and construct a retrieval coding library.
[0224] S5, training feature extraction network and compact coding mapping unit. Semantic similarity weighted center loss, quantization loss and hash coding loss are used to constrain model training. The specific constraint design refers to the specific description in the previous article and will not be repeated here. The back propagation algorithm is used to update and optimize the network parameter weights until the model loss area converges. In this example, image restoration model training and evaluation are completed on the Pytorch platform. The model is trained on a single GeForce RTX3090Ti GPU (24GB) and the batch size is set to 16. The learning rate used is 2.5×10 -4The SGD optimizer optimizes the network, and the weight decay coefficient and momentum coefficient are set to 1×10 -4 and 0.91, and the learning rate is reduced by 10% after the first iteration. For datasets with less than 20,000 images in the training set, 40 epochs of 30 iterations per epoch are trained, while for other datasets, 50 epochs of 30 iterations per epoch are trained. The results of the best performing model on the test set are finally reported.
[0225] S6, after the model training is completed, fine-grained image retrieval is performed, and the test data set of each data set is input into the trained model, and the model output result is the retrieval result. This example uses mAP as the final evaluation indicator, and the encoding lengths are 12 bits, 24 bits, 32 bits, and 48 bits respectively. For the CUB200-2011 dataset, the retrieval accuracies under the four length encodings are 83.76%, 88.92%, 89.37% and 90.28% respectively; for the Aircraft dataset, the retrieval accuracies under the four length encodings are 75.32%, 85.91%, 86.47% and 86.99% respectively; for the Food101 dataset, the retrieval accuracies under the four length encodings are 85.22%, 88.97%, 89.11% and 89.96% respectively; for the NABirds dataset, the retrieval accuracies under the four length encodings are 67.83%, 79.61%, 82.85% and 85.27% respectively; for the VegFru dataset, the retrieval accuracies under the four length encodings are 88.70%, 92.95%, 93.86% and 93.27% respectively. The results show that the present invention can effectively complete fine-grained image retrieval tasks and performs well under multiple length encodings.
[0226] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.
[0227] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0228] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A fine-grained image retrieval method based on compact representation modeling and semantic label guidance, characterized by: The following steps are involved: Step S1, preparing an image dataset containing fine-grained categories, performing data preprocessing and data division, and obtaining a preprocessed image dataset; Step S2, constructing a feature extraction network, including constructing a feature extraction network based on a multi-layer visual Transformer as a backbone network, for extracting visual features of images in the preprocessed image data set to obtain visual word vectors, the feature extraction network also includes a global level feature learning branch and a local level feature learning branch, wherein the local level feature learning branch includes a MUSTA module; Step S3, generating a hash coding center, building a bag-of-words model based on the fine-grained image labels, and counting the frequency of each word appearing in the category label as the word frequency, calculating the semantic similarity matrix between all category labels based on the statistically obtained word frequency, and finally generating the coding center of each category with the semantic similarity matrix as the weight; Step S4, constructing a compact coding mapping unit, inputting the visual word vector extracted by the feature extraction network into the constructed coding mapping unit, and using the activation function in the coding mapping unit to map the high-dimensional real-valued features of the visual word vector into low-dimensional binary codes, to obtain a compact representation of the image, and constructing a retrieval coding library with the obtained compact representation of the image; Step S5, training the constructed feature extraction network and compact coding mapping unit to obtain trained feature extraction network and compact coding mapping unit; Step S6, perform fine-grained image retrieval, input the image to be detected into the trained feature extraction network and the compact coding mapping unit to obtain the high-dimensional features of the image to be detected, map the high-dimensional features of the image to be detected into low-dimensional codes as the codes to be detected, match the obtained codes to be detected with the codes in the retrieval code library, and the sample with the smallest distance is the final retrieval result.
2. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to claim 1 is characterized in that: The step S2 specifically includes the following steps: Step S2.1, constructing a feature extraction network, first constructing a multi-layer visual Transformer including a global level feature learning branch and a local level feature learning branch as the backbone network of the feature extraction network, wherein the local level feature learning branch includes a MUSTA module; Step S2.2, input the preprocessed image dataset and use the constructed multi-layer visual TransformerΦ ViT (·) Extract a sequence of global-level visual word vectors from image I in the input preprocessed image dataset: T gl =Φ ViT (I); Step S2.3, in parallel, the set MUSTA module is used to select several relatively significant visual word vectors from each attention head of each Transformer layer in all Transformer layers except the last layer of the multi-layer visual Transformer, and all the significant visual word vectors selected from each Transformer layer except the last layer are aggregated into the final input sequence Z′ L-1 , where L represents the total number of Transformer layers included in the multi-layer visual Transformer; Step S2.4: The final input sequence Z′ after aggregation L-1 Input into the last Transformer layer in the local level feature learning branch, and after processing, the local level visual word vector sequence is obtained:
3. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to claim 2 is characterized in that: The step S2.3 specifically includes the following steps: Step S2.3.1, the total number of Transformer layers included in the multi-layer visual Transformer is L layers, and the input of the last Transformer layer in the local level feature learning branch is Z′ L-1 , the attention weights of each layer of the first L-1 layers of Tranformer with M attention heads are expressed as: Among them, A l represents the attention weight of the lth Transformer layer, l represents the current Transformer layer number and 1≤l≤L-1, Indicates A l The attention weight of the mth attention head in , express N T The attention weight corresponding to the visual word vector, m represents the current attention head number in the Tranformer layer, N T Represents the number of non-category visual word vectors for each attention head; Step S2.3.2, in the MUSTA module, the attention weights of the Lth Tranformer layer and before are multiplied layer by layer to obtain the fused attention weight used to record the forward propagation of the visual word vector from the input layer to the higher layer: Among them, a′ represents the fusion attention weight, w represents the number of the Transformer layer, and A w represents the attention weight of the w-th Tranformer layer and 1≤w≤l; Step S2.3.3, select the attention weight corresponding to the category visual word vector from the fused attention weight a′: And filter out the indexes corresponding to the n' visual word vectors with the highest weight values from each attention head: Among them, v l represents the fused attention weight corresponding to the category visual word vector in the lth layer Transformer layer, Indicates v l The attention weight of the mth attention head in , K l Represents the index values corresponding to several visual word vectors with the highest fused attention weights in the lth layer Transformer layer, K l The visual word vector index value of the mth attention head in . The TopSelection(·) function means using the max(·) function to filter the index corresponding to the highest attention weight. Step S2.3.4, in the output sequence of each Transformer layer of the first L-1 layers of Transformer layers, find the index value K that matches the selected index value of that layer l The corresponding visual word vector: T s,l =T l [K l ] Among them, T l Represents the output sequence of the lth layer Transformer layer; Step S2.3.4, aggregate the visual word vectors of the output sequences of each Transformer layer of the first L-1 layers of Transformer layers obtained above with the category visual word vectors of the original input sequence of the Lth Transformer layer through the MUSTA module as the final input sequence of the last Transformer layer in the local level feature learning branch: in, Represents the category visual word vector of the original input sequence of the L-th Transformer layer.
4. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to claim 3 is characterized in that: The step S3 specifically comprises the following steps: Step S3.1, generate a set of q binary centers with at least distance d between K subcategories: Where k represents the fine-grained category number, c k represents the hash code center of the kth class, d is the minimum distance between any two representation centers obtained based on the Gilbert-Varshamov boundary theory, K represents the total number of fine-grained categories, and q represents the hash code length; Step S3.2, construct a category semantic similarity matrix, adopt a bag-of-words model containing w' words, and represent the K fine-grained labels as a feature matrix V based on word occurrence probability. The similarity matrix is expressed as: S=V·V T ; Based on the category similarity matrix S, the above minimum distance d is replaced by a fine-grained weighted distance constraint d between two categories. i,j : d i,j =(1+S i,j )·d; Then, the hash code centers of each fine-grained category are generated by solving the following optimization problem: s.t.||c i -c j || H ≥d i,j ,1≤i,j≤K,i≠j, in,‖·‖ H is the Hamming distance, i and j represent the fine-grained category numbers, and c K 、c i 、c j Respectively represent the hash code center of the Kth, ith, and jth categories, d i,j represents the fine-grained weighted distance constraint between the i-th class and the j-th class, S i,j It represents the similarity between the i-th and j-th labels.
5. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to claim 4 is characterized in that: The step S4 specifically comprises the following steps: Step S4.1, by respectively embedding the visual word vector sequence T at the global level gl and the local level visual word vector sequence T loc Apply global average pooling to obtain the global level real-valued feature x gl and local level real-valued features x loc ; Step S4.2, use the global level linear mapping matrix W gl and the local level linear mapping matrix W loc , respectively, the high-dimensional global level real-valued features x gl and high-dimensional local-level real-valued features x loc Mapped to low-dimensional global level real-valued features h gl and low-dimensional local level real-valued features h loc ; Step S4.3, transform the low-dimensional global level real-valued feature h gl and low-dimensional local level real-valued features h loc Mapping to low-dimensional compact encoding b = {b gl ,b loc }, we get a compact representation of the image, where b gl represents the global level low-dimensional compact encoding, b loc Represents a local-level low-dimensional compact encoding. Step S4.4, the low-dimensional compact codes corresponding to the samples of the images in all the input preprocessed image data sets are stored according to the image sample numbers in each gallery, and a retrieval code library with an index set of Ω is constructed.
6. The fine-grained image retrieval method based on compact representation modeling and semantic label guidance according to claim 5, characterized in that: The step S5 specifically includes the following steps: Step S5.1, embed the similarity matrix S into the compact code b of the nth sample of the kth class n and the corresponding k-th hash code center c k , we get the similarity between the compact code and its corresponding hash code center: Among them, Sim(·,·) represents the scaled cosine similarity calculation, P n,k Represents the compact code b n and the corresponding k-th hash code center c k similarity, t represents the fine-grained category number, K represents the total number of fine-grained categories, S k,t Indicates the similarity between the k-th and t-th class labels, c t represents the hash code center of the tth class; Step S5.2, establish the semantic similarity weighted center loss function: Among them, Z n,k is an indicator factor. If the nth input image I n The compact encoding of n belongs to the kth category, then Z n,k =1, otherwise Z n,k =0; Step S5.3, establish quantization loss and hash coding loss function: Among them, λ1 and λ2 represent the loss term coefficients, u represents the image sample number in the gallery, v represents the query sample number, Ω represents the index set of gallery samples in the training data, Γ represents the index set of query samples in the training data, and b u , b v Represents the low-dimensional hash code of samples u and v, respectively, u represents the low-dimensional real-valued encoding of sample u, L uv represents paired sample supervision information and L uv ∈{-1,+1} (Q×D) , Q represents the number of query samples, D represents the number of gallery samples, The first term is to convert the real-valued low-dimensional feature vector h u The second term is the general semantic preservation loss. Step S5.4, the quantization loss and hash coding loss function Similarity with semantic weighted center loss function Combined, we finally get the total loss function that optimizes the entire network: Step S5.5: The final total loss function includes the losses of the feature extraction network and the compact coding mapping unit. During the gradient back-propagation process, the network parameters of the feature extraction network and the compact coding mapping unit are updated at the same time until the loss of the loss function converges, thus obtaining a trained feature extraction network and a compact coding mapping unit.
Citation Information
Patent Citations
Deep hash image retrieval method fusing semantic information and multi-level similarity
CN109977250A
Fine-grained image-text retrieval method and system based on Transform model
CN114780766A
Dressing pedestrian re-identification method based on multilayer dynamic concentration and local pyramid aggregation
CN117635973A
Long sequence modeling via state space model (SSM)-enhanced transformer
US20240202583A1