Deep hash image retrieval method based on feature fusion and dynamic converter
By combining the feature fusion and dynamic transformer of CNN and ViT, adopting dynamic convolution and LoRA modules, and designing a multi-branch loss function, the problem of insufficient local and global feature extraction in deep hashing methods is solved, the robustness and retrieval performance of hash codes are improved, and it is suitable for large-scale image retrieval.
Patent Information
- Application Number
- CN202510862402.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-03
AI Technical Summary
Existing deep hashing methods based on CNN and ViT have problems in image retrieval, such as insufficient extraction of local and global features, high computational complexity, and poor robustness of hash codes. In addition, the fixed loss function leads to imbalanced samples of similar images and large hash code errors.
Combining the local feature extraction capability of CNN with the global modeling advantage of ViT, dynamic convolution, KAN module and LoRA low-rank adaptation mechanism are introduced, and a jointly optimized multi-branch loss function is designed. The semantic consistency and retrieval performance of hash codes are improved through feature fusion and dynamic transformer.
It improves the efficiency and accuracy of image retrieval, reduces computational complexity, makes the model suitable for resource-constrained environments, and achieves efficient and accurate large-scale image retrieval.
Smart Images

Figure CN120744154A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and computer vision technology, and in particular to a deep hash image retrieval method based on feature fusion and dynamic transformer. Background Art
[0002] With the popularization of the Internet and mobile devices, the amount of image data has exploded. How to achieve fast and accurate retrieval in massive image data has become an urgent problem to be solved in the current image retrieval field. Traditional image retrieval methods rely on manually designed low-dimensional features, such as color histograms and scale-invariant feature transforms (SIFTs). These methods have obvious shortcomings in describing the deep semantics of images and have low retrieval efficiency in large-scale data, making it difficult to meet the application's requirements for real-time and accuracy. In recent years, the rise of deep learning technology has greatly promoted the development of image retrieval technology, especially deep hashing methods have attracted more attention. Deep hashing methods map high-dimensional image data into compact binary hash codes by training deep neural networks. This not only effectively maintains the semantic similarity of images, but also greatly reduces storage space and computational complexity, significantly improving the efficiency and accuracy of large-scale image retrieval. Deep hashing methods are gradually becoming a key technology for solving the challenges of large-scale image retrieval, and are widely used in social media, e-commerce, intelligent monitoring and other fields, promoting the application and development of image retrieval systems.
[0003] At present, the relevant technologies are mainly divided into CNN-based deep hashing methods and ViT-based deep hashing methods. CNN has more advantages in extracting local features of images, while ViT models are gradually becoming more popular in image tasks due to their global feature extraction capabilities. However, the single structure has the following problems: (1) CNN models focus more on local features but ignore global relationships; (2) Although ViT models have the advantage of focusing on global semantics, they do not capture local features of images sufficiently and have high computational complexity. In addition, most current hash networks use fixed loss functions, which have problems such as imbalanced similar image samples and large errors caused by discrete hash codes, resulting in poor robustness and insufficient discrimination ability of the generated hash codes.
[0004] Therefore, this field urgently needs a deep hash image retrieval method based on feature fusion and dynamic transformer to solve the above problems. Summary of the Invention
[0005] In order to solve or partially solve the problems existing in the related technologies, this application provides a deep hash image retrieval method based on feature fusion and dynamic transformer. By combining the local feature extraction capability of CNN with the global modeling advantages of ViT, dynamic convolution, KAN module and LoRA low-rank adaptation mechanism are introduced, and a jointly optimized multi-branch loss function is designed to effectively improve the semantic consistency and retrieval performance of the hash code, and realize efficient and accurate large-scale image retrieval.
[0006] The first aspect of the present application provides a deep hash image retrieval method based on feature fusion and dynamic transformer, comprising the following steps:
[0007] Step 1: Obtain image dataset and establish semantic association relationships between samples;
[0008] Step 2: extracting image features through a dynamic transformation network that fuses local and global features, wherein the dynamic transformation network includes an adaptive feature extraction module and a global semantic modeling module;
[0009] Step 3: Compress the extracted image features into compact binary hash codes and compare the similarities between images.
[0010] A second aspect of the present application provides an electronic device, including:
[0011] processor; and
[0012] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.
[0013] The technical solution provided by this application may have the following beneficial effects: the deep hash image retrieval method based on feature fusion and dynamic transformer proposed in this application has a clear overall process and is easy to implement. Feature extraction and hash code generation are completed in a unified network framework and can be deployed in various image retrieval systems; the proposed CNN and ViT fusion network, combined with dynamic convolution and KAN modules, can fully explore the local and global features of the image and improve the feature expression ability;
[0014] In addition, this application reduces computational complexity by introducing the LoRA module, making the model more suitable for resource-constrained environments; the designed multi-branch joint loss function combines hash loss, metric alignment and quantization constraints to comprehensively optimize the hash learning process and effectively improve retrieval accuracy and efficiency; it performs well on large-scale image datasets and has good application promotion value.
[0015] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0017] Figure 1 1 is a flow chart of a deep hash image retrieval method based on feature fusion and dynamic transformer according to an embodiment of the present application;
[0018] Figure 2 The precision-recall curves of hash codes of different lengths for the deep hash image retrieval method based on feature fusion and dynamic transformer shown in the embodiment of the present application;
[0019] Figure 3 This is the accuracy of the deep hash image retrieval method based on feature fusion and dynamic transformer shown in the embodiment of the present application when hash codes of different lengths are used as search indexes.
[0020] Figure 4 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0022] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0023] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0024] To address the above problems, an embodiment of the present application provides a deep hash image retrieval method based on feature fusion and dynamic transformer, which can improve the semantic consistency and retrieval performance of hash codes and achieve efficient and accurate large-scale image retrieval.
[0025] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0026] Example 1:
[0027] See also Figure 1-3 , a deep hash image retrieval method based on feature fusion and dynamic transformer, comprising the following steps:
[0028] Step 1: Data preprocessing: Obtain image datasets and establish semantic associations between samples;
[0029] The data preprocessing process includes data preparation and similarity matrix construction, which structures the original image data and clarifies the relationship between images, providing data support for subsequent image retrieval training.
[0030] By collecting a collection of images and annotating them with category labels, we use one-hot encoding to represent them; we construct a similarity matrix to determine the relationship between positive and negative sample pairs, where positive sample pairs belong to the same category and negative sample pairs do not belong to the same category;
[0031] Specifically, when preparing data and constructing the similarity matrix, the specific steps include:
[0032] S11, collect image dataset: prepare a set of N images {x1,x2,...,x N}, and annotate each image with the corresponding category label y i , represented by one-hot encoding;
[0033] S12, construct the similarity matrix S:
[0034]
[0035] S13, establish sample pair set: construct positive sample pair set S respectively + (S ij =1, that is, image x i with x j Similar, that is, y i =y j , belong to the same category) and the negative sample set S - (S ij =0, that is, image x i with x j Not similar, that is, y i ≠y j, do not belong to the same class), which is used for subsequent loss function weighting processing.
[0036] Step 2: Feature fusion extraction: Extract image features through a dynamic transformation network that fuses local and global features. The dynamic transformation network includes an adaptive feature extraction module and a global semantic modeling module.
[0037] The feature fusion extraction process includes the construction of a feature extraction network (CNN+ViT fusion), using two improved feature extraction network modules for feature extraction, fusing local and global image information to obtain a more discriminative image feature representation; including a local feature extraction module and a global feature extraction module;
[0038] By setting up a multi-scale dynamic convolution module, convolution kernels of different scales are weighted and fused through the attention weight vector, and convolution and activation functions are performed to extract local features. The local features are mapped to the embedding space, and after adding learnable classification tags and position encoding, they are input into the global semantic modeling module, and finally the global feature vector corresponding to the classification tag is output.
[0039] Furthermore, the global semantic modeling module is a Transformer encoder. Each encoder layer includes layer normalization, a multi-head self-attention mechanism, and a nonlinear mapping module. The nonlinear mapping module uses a KAN module to replace the traditional MLP. The KAN module uses a combination of multiple single-variable nonlinear functions to represent the nonlinear mapping. The Transformer encoder uses the LoRA mechanism to perform low-rank adaptive optimization, reducing trainable parameters and computational overhead, while enabling rapid model adjustment.
[0040] Specifically, the feature fusion extraction process includes the following steps:
[0041] S21, set K groups of convolution kernels of different scales {W1,W2,...,W K To build a dynamic convolution module, different scales correspond to different receptive fields, which can enhance the multi-level perception ability of the image. This embodiment uses three convolution kernels of different scales (3×3, 5×5, and 7×7) to better extract multi-dimensional image features from edge texture to the whole image.
[0042] S22, calculate the attention weight vector for the input image Where FC(x) = W fc x+b fc is the output of x after the fully connected layer (FC), is the fully connected weight matrix, d is the dimension of the input vector, is the bias vector, Represents a K-dimensional real number space, that is, the weight vector ω is a K-dimensional real number vector;
[0043] S23, according to the dynamic weight ω of the i-th convolution kernel i and its parameter W i , weighted fusion convolution kernel forms a dynamic convolution kernel:
[0044] S24, perform convolution and activation functions to extract local features:
[0045]
[0046] Where f(·) is the activation function;
[0047] S25, since the feature dimension of CNN output is not consistent with the input dimension of ViT, the convolution feature y i Mapping to the embedding space to satisfy the input of ViT:
[0048] z i =y i W emb +b emb
[0049] where y i is the feature vector of the i-th image output by CNN, z i The feature vector of the input ViT for the i-th image, is the embedding weight matrix, is the embedding bias vector;
[0050] S26 adds learnable classification tags (cls-token) and position encodings (pos-embedding). Cls-token carries global semantic information and is trained by the network model and continuously updated during training. Position encoding is the "spatial position information" added to ViT to prevent the model from confusing the position order. Each position corresponds to a feature vector, which is also used as a model parameter in training after initialization:
[0051] Z input =z cls ,z1+e1,...,z n +e n
[0052] where Z input Represents the complete input sequence of the encoder, which is a set of vectors, z cls Represents cls-token;
[0053] S27, input stacked 12-layer Kansformer encoder:
[0054] Z output =T(Z input )
[0055] Z output Z input The output after T, T is the encoder containing MSHA and KAN modules. The internal parameters of each KAN layer are optimized by low-rank adaptation through the LoRA mechanism. The following ③④ are all optimizations of the KAN part in the encoder in this invention;
[0056] Each layer includes:
[0057] ① Layer normalization (LayerNorm) makes the output distribution of each layer more stable to improve training stability and accelerate convergence;
[0058] ② Multi-head self-attention mechanism (MHSA), which enables the model to focus on different semantic relations in multiple subspaces in parallel;
[0059] ③ KAN replaces traditional MLP. MLP is essentially a linear plus nonlinear superposition, with limited expressive power. KAN, on the other hand, uses a combination of multiple (D in this embodiment) single-variable nonlinear functions φ1(x) to better represent nonlinear mappings. Furthermore, with the same fitting accuracy, KAN requires far fewer parameters than MLP and can better focus on complex information in the image.
[0060] KAN(x)=[φ1(x),...,φ D (x)]
[0061] ④Introducing the LoRA module to adapt the KAN parameters to low rank:
[0062]
[0063] Where W KAN Represents the original weight matrix and does not need to be retrained; A is a D×r learnable low-rank matrix, and B is an r×D learnable low-rank matrix. During training, only A and B are trained, which can greatly reduce the trainable parameters and computational overhead, and also allows for rapid model adjustment;
[0064] S28, finally output the feature vector v corresponding to the classification label cls , used to generate hash codes;
[0065]
[0066] in, Represents the vector with index 0 in the encoder output sequence, which is the updated value of cls-token.
[0067] Step 3: Hash code generation: compress the extracted image features into a compact binary hash code to achieve efficient image representation;
[0068] The hash code generation process includes converting the image feature vector v extracted above intocls Compressed into hash codes for fast image matching; the core of this step is to use a set of compact "feature fingerprints" containing only 0s and 1s to quickly compare the similarities between images, greatly improving image retrieval efficiency;
[0069] Specifically, the hash code generation process includes:
[0070] S31, projects the global feature vector into the hash space through linear mapping;
[0071] h i =W hash v cls +b hash
[0072] Among them, h i Represents the hash chain, which is the result of mapping features to hash space, W hash v cls Represents the hash mapping weight matrix, through which v cls Projected into hash space, b hash Represents the hash map bias item, which provides bias compensation for the hash output;
[0073] S32, uses continuous relaxation approximate binarization in the training phase;
[0074] b i =tanh(h i );
[0075] S33, in the inference stage, a hard binarization method is used to generate a 0 / 1 hash code, which is the binary hash code ultimately used for image retrieval;
[0076]
[0077] It should be noted that this embodiment adopts different binarization methods in the training stage and the inference stage; in the training stage, the binarization operation (such as the sign function sign(x)) is discrete and non-differentiable at point 0. Direct use for training will result in the inability to transfer gradients during back propagation and the inability to optimize the model parameters. Therefore, in the training stage, a continuously differentiable function (such as tanh(x)) is needed to approximate the binarization process so that the hash code output is a continuous value close to ±1, thereby allowing gradient back propagation and driving the model to learn the correct feature mapping parameters; in the inference stage, the final hash code is required to be in a strict 0 / 1 binary form to significantly reduce storage overhead and retrieval calculation complexity; hard binarization maps continuous features into discrete binary codes, which are directly used for image retrieval.
[0078] Step 4: Joint loss optimization: Design a multi-branch loss function to optimize the hash encoding process to improve the semantic consistency and retrieval performance of the hash code;
[0079] The joint loss optimization process involves designing and optimizing a multi-branch joint loss function. To guide the model to correctly learn hash codes, this embodiment designs a triple loss structure (weighted contrast loss, metric alignment loss, and quantization loss) and integrates them into a total loss function for optimization.
[0080] Furthermore, the multi-branch loss function includes: calculating the weighted negative log-likelihood loss to balance the weights of positive and negative sample pairs; calculating the metric alignment loss to constrain the hash code distance of similar images; calculating the quantization loss to reduce the error of the binarization process; and fusing the weighted negative log-likelihood loss, metric alignment loss, and quantization loss to obtain the total loss function.
[0081] Specifically, the process of joint loss optimization includes:
[0082] S41, calculate the weighted negative log-likelihood loss (HashLoss):
[0083]
[0084] where a ij Represents the sample pair weighting factor, which represents the image pair (x i ,x j ) is used to adjust the proportional influence of positive and negative samples in the loss. The subscript T represents the vector transpose, that is, represents the inner product between two hash vectors;
[0085] Since the number of positive and negative sample pairs in the training set is quite different, it is easy to cause the loss function to be biased towards the negative sample pairs with a larger proportion, so that the model cannot ensure the feature alignment of similar images. Therefore, the weighting factor a is set. ij , to ensure that positive and negative sample pairs have equal normalized weights in the total loss and reduce the problem of class imbalance:
[0086] S42, metric alignment loss (MetricLoss):
[0087]
[0088] Where k is the dot product value of the hash code of the similar image pair, which is equal to the length of the hash code;
[0089] S43, in order to reduce the error between the continuous hash code in the training phase and the discrete hash code in the inference phase, this embodiment introduces quantization loss (QuantizationLoss):
[0090]
[0091] The quantization loss is used to make the real-valued hash code b i Try to be as close to ±1 as possible, thereby reducing the error in the conversion to binary hash code and improving the stability of the retrieval system;
[0092] S44, after completing the design of the three losses, merge them into the total loss function:
[0093] L total =αL NLL +βL metric +γL quant
[0094] Among them, α, β, and γ are hyperparameter weights, and backpropagation is used to update the network weights (LoRA only optimizes the low-rank matrices A and B, and the rest of the parameters are frozen).
[0095] Furthermore, the process of improving the semantic consistency and retrieval performance of hash codes includes:
[0096] Taking the total loss function as the optimization target, the model parameters are gradient-calculated using the standard back-propagation algorithm, and the Adam optimizer is used for iterative updates. This process is repeated until the total loss function converges or the preset number of training rounds is reached, thus completing the training.
[0097] Specifically, forward propagation generates prediction results, the three losses are integrated to calculate the total loss value, and then the trainable parameters in the network (dynamic convolution kernel weights, KAN connection structure, LoRA low-rank matrix) are reversely updated; the above process is executed repeatedly until the loss function converges or the preset number of training rounds is reached, thereby completing the network training.
[0098] See also Figure 1-3 , Figure 2 This is the Precision-Recall curve of the method in this embodiment, which evaluates the performance under different bit sizes (16 bits, 32 bits, 48 bits, and 64 bits). High bit sizes (such as 64 bits) always maintain high Precision, indicating that the model can effectively identify relevant results and reduce false positives during image retrieval, with guaranteed accuracy. The retrieval quality is especially outstanding in the early stages of retrieval (small index) and in scenarios with high precision requirements (high bit settings).
[0099] Figure 3 This is the Precision-image search index curve of the method in this embodiment. The performance is also evaluated under different bit sizes (16 bits, 32 bits, 48 bits, and 64 bits).
[0100] At different bit rates, the variation of Precision with image search index or Recall follows a consistent pattern and the curve is smooth, indicating that the model performs stably during the retrieval process. It does not experience significant fluctuations in precision due to expansion of the search scope (increase in index) or increased recall requirements, and can reliably output retrieval results.
[0101] The deep hash image retrieval method based on feature fusion and dynamic transformer proposed in this embodiment has a clear overall process and is easy to implement. Feature extraction and hash code generation are completed in a unified network framework and can be deployed in various image retrieval systems. The proposed CNN and ViT fusion network, combined with dynamic convolution and KAN modules, can fully explore the local and global features of the image and improve the feature expression ability. By introducing the LoRA module, the computational complexity is reduced, making the model more suitable for resource-constrained environments. The designed multi-branch joint loss function combines hash loss, metric alignment and quantization constraints to comprehensively optimize the hash learning process and effectively improve retrieval accuracy and efficiency. This method performs excellently on large-scale image datasets and has good application and promotion value.
[0102] Example 2:
[0103] See also Figure 4 , the electronic device 1000 includes a memory 1010 and a processor 1020.
[0104] The processor 1020 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0105] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage may be a readable and writable storage device. The permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (e.g., a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0106] The memory 1010 stores executable codes. When the executable codes are processed by the processor 1020 , the processor 1020 may execute part or all of the above-mentioned methods.
[0107] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0108] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium) on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.
[0109] Example 3:
[0110] This example demonstrates the beneficial effects of this method by establishing the following experiments:
[0111] 1. Experimental conditions and performance evaluation standards:
[0112] The performance experiments of this invention were run on a server equipped with NVIDIA RTX 4060 GPU, the optimizer used was Adam, and the learning rate was set to 8×10 -5 Experiments were conducted on the CIFAR-10 and NUS-WIDE datasets to compare and evaluate the proposed method with existing deep hashing methods. The advantages can be further illustrated by the following simulation experiments:
[0113] The experimental environment is:
[0114] (1) Hardware environment: AMD Ryzen 59600X 6-Core Processor, 3.90GHz, 32G memory;
[0115] (2) The software environment is: Windows 11 operating system, using PyCharm development platform and Python programming language.
[0116] 2. Experimental content
[0117] In related technologies, Xia et al. (“Xia R, Pan Y, Lai H, et al. Supervised hashing for image retrieval via image representation learning[C] / / Proceedings of the AAAI conference on artificial intelligence. 2014, 28(1).”) first proposed combining CNN with hash learning to construct a deep hashing algorithm, which greatly improved retrieval performance. However, this method lacks an end-to-end framework, is inconvenient to use, and takes a long time. Cao et al. (“Cao Z, Long M, Wang J, et al. Hashnet: Deep learning to hash by continuation[C] / / Proceedings of the IEEE international conference on computer vision. 2017: 5608-5617.”) addressed the non-differentiable problem of hash codes in deep hashing methods by using the tanh function to implement backpropagation, successfully solving the non-differentiable problem and improving the efficiency of deep hashing construction. Xu et al. (“Xu C, Chai Z, Xu Z, et al. HHF: Hashing-guided hinge function for deep hashing retrieval[J]. IEEE Transactions on Multimedia, 2022, 25: 7428-7440.”) used a hinge loss function to reduce the ambiguity of semantic information and the generation of incorrect hash codes during hashing learning, effectively solving the problem of semantic degradation. Ren et al. (“Ren H, Cheng S, Wang L, et al. Convolution-based Visual Transformer with Dual Space Shifting Attention hashing for image retrieval[J]. AppliedSoft Computing, 2025, 171: 112752.”) developed a ViT-based hashing method that leverages the latest advances in deep learning and the Transformer architecture to improve retrieval performance.Li et al. (“Li T, Zhang Z, Pei L, et al. HashFormer: Vision transformer based deep hashing for image retrieval[J]. IEEE Signal Processing Letters, 2022, 29: 827-831.”) introduced a novel average precision loss based on ViT to directly optimize retrieval performance. Yu et al. (“Yu J, Qiu H, Chen D, et al. Weighted contrative hashing[C] / / Proceedings of the Asian Conference on Computer Vision. 2022: 3861-3876.”) proposed a weighted contrastive hashing model with a mutual attention mechanism to enhance semantic interactions between image patches, alleviate the inconsistency problem of anchor point augmentation images, and improve hash code learning through similarity-weighted training. Wang et al. (“Wang W, Guo Z, Jiang W, et al. CrossHash: Cross-scale Vision Transformer Hashing for Image Retrieval[C] / / ICASSP 2025-2025IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2025: 1-5.”) proposed a deep hashing method based on a cross-scale Transformer to effectively extract and fuse multi-scale features, overcoming the limitations of traditional ViT's single-scale approach. Liu et al. (“Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C] / / Proceedings of the IEEE / CVFinternational conference on computer vision. 2021: 10012-10022.”) proposed the SwinTransformer, which uses a local window attention mechanism, a sliding mechanism, and a layered design to achieve a balance between computational efficiency and accuracy. This method combines the local feature extraction capability of CNN with the global context modeling of Transformer, and performs excellently, making it suitable for processing multi-scale and complex images.Zhang et al. (“Zhang B, Zhang Y. MSCViT: ASmall-size ViT architecture with Multi-Scale Self-Attention Mechanism for Tiny Datasets[J]. arXiv preprint arXiv:2501.06040, 2025.”) proposed a lightweight ViT model that uses an asymmetric learning strategy to effectively capture local and global features, solving the problems of insufficient global information and limited model generalization ability.
[0118] Experiment 1: Ablation experiment
[0119] Performance comparison of different models under different data sets
[0120] The model proposed in this embodiment combines dynamic convolution with improved ViT for deep hash image retrieval, and uses ablation experiments to compare the model performance. The comparison results are shown in Table 1.
[0121] Table 1. Comparison of mAP in ablation experiments
[0122]
[0123] Table 1 shows the results of ablation experiments. The introduction of the KAN module and LoRA fine-tuning significantly improves feature extraction performance, while the overall design of the hash loss function further enhances the quality of the hash codes. Specifically, the metric loss aligns the inner product of the hash codes with the target similarity, ensuring a smaller Hamming distance between similar images and a larger distance between dissimilar images. Furthermore, the quantization loss effectively reduces the error in the conversion from continuous values to binary hash codes. These results demonstrate the importance of each component of this method and their synergistic effect.
[0124] Experiment 2: Model Complexity Analysis
[0125] The deep hashing image retrieval method based on feature fusion and dynamic transformer proposed in this paper is compared with the model complexity of Li method (“Li T, Zhang Z, Pei L, et al. HashFormer: Vision transformer based deephashing for image retrieval[J]. IEEE Signal Processing Letters, 2022, 29: 827-831.”) and Ren method (“Ren H, Guo J, Cheng S, et al. Pooling-based Visual Transformerwith low complexity attention hashing for image retrieval[J]. Expert Systemswith Applications, 2024, 241: 122745.”) in tabular form. The comparison results are shown in Table 2.
[0126] Table 2 Comparison of model complexity of different methods
[0127]
[0128] As can be seen in Table 2, CVIH benefits from the LoRA fine-tuning technique, significantly reducing the training parameters of the Kansformer encoder. Compared with the Li method (which requires fine-tuning all parameters, approximately 22.5 million), the method of this embodiment has only 5.8 million trainable parameters, reducing parameter overhead by approximately 74%. Compared with the Ren method, the trainable parameters of this embodiment are also reduced by approximately 65%, which is mainly due to the low-rank decomposition design of LoRA. On the CIFAR-10 dataset, the average single-image inference time of the method of this embodiment is 12.3 milliseconds, which is basically the same as the Li method (11.9 milliseconds) and the Ren method (12.1 milliseconds). The results show that despite the introduction of the KAN module and the LoRA fine-tuning mechanism, the proposed method does not significantly increase the computational overhead and is suitable for real-time image retrieval tasks.
[0129] Experiment 3: Comparison of mAP of different methods
[0130] The deep hash image retrieval method based on feature fusion and dynamic transformer proposed in this invention is compared with the HashFormer algorithm and PTLCH algorithm in the Li method (“Li T, Zhang Z, Pei L, et al. HashFormer: Vision transformer based deephashing for image retrieval[J]. IEEE Signal Processing Letters, 2022, 29: 827-831.”) and the Ren method (“Ren H, Guo J, Cheng S, et al. Pooling-based Visual Transformerwith low complexity attention hashing for image retrieval[J]. Expert Systemswith Applications, 2024, 241: 122745.”) in terms of entropy in tabular form. The comparison results are shown in Table 3.
[0131] Table 3 Comparison of mAP of different algorithms
[0132]
[0133] As shown in Table 3, on the CIFAR-10 dataset, the method in this embodiment outperforms PTLCH by 2.1%, 1.6%, 1.9%, and 2.3% under four different hash code lengths, respectively. On the NUS-WIDE dataset, it outperforms HashFormer by 3.7%, 4%, 4.3%, and 4.8%, respectively, indicating that the method in this paper consistently outperforms other compared methods in overall performance.
[0134] Through experimental comparative analysis, the method of this embodiment has better model performance while reducing time overhead as much as possible, and can effectively improve retrieval accuracy. The present invention improves the feature expression ability in image retrieval, improves image retrieval accuracy and efficiency, and further improves the performance of large-scale image retrieval without increasing computational overhead.
[0135] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A deep hash image retrieval method based on feature fusion and dynamic transformer, characterized in that: The following steps are involved: Step 1: Obtain image dataset and establish semantic association relationships between samples; Step 2: extracting image features through a dynamic transformation network that fuses local and global features, wherein the dynamic transformation network includes an adaptive feature extraction module and a global semantic modeling module; Step 3: Compress the extracted image features into compact binary hash codes and compare the similarities between images.
2. A deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 1, characterized in that: Also includes: Step 4: Design a multi-branch loss function to optimize the hash encoding process.
3. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 1 is characterized in that: The process of obtaining an image dataset and establishing semantic association relationships between samples includes: Collect image collections and annotate them with category labels, using one-hot encoding. A similarity matrix is constructed to determine the relationship between positive and negative sample pairs between samples, where positive sample pairs belong to the same category and negative sample pairs do not belong to the same category.
4. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 1 is characterized in that: The process of extracting image features by a dynamic transformation network fusing local and global features includes: Set up a multi-scale dynamic convolution module, fuse convolution kernels of different scales through the weighted attention weight vector, and perform convolution and activation functions to extract local features; The local features are mapped to the embedding space, and after adding learnable classification tags and position encoding, they are input into the global semantic modeling module, and the global feature vector corresponding to the classification tags is output.
5. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 1 is characterized in that: The global semantic modeling module is a Transformer encoder, and each layer of the encoder includes layer normalization, a multi-head self-attention mechanism and a nonlinear mapping module. The nonlinear mapping module uses a KAN module to replace the traditional MLP. The KAN module uses a combination of multiple single-variable nonlinear functions to represent the nonlinear mapping.
6. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 5 is characterized in that: The Transformer encoder performs low-rank adaptation optimization through the LoRA mechanism.
7. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 4 is characterized in that: The process of generating the hash code includes: Project the global feature vector to the hash space through linear mapping; During the training phase, a continuous relaxation approximate binarization method is used; In the inference stage, hard binarization is used to generate 0 / 1 hash codes, which are finally used as binary hash codes for image retrieval.
8. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 2 is characterized in that: In step 4, the multi-branch loss function includes: Calculate the weighted negative log-likelihood loss to balance the weights of positive and negative sample pairs; Compute a metric alignment loss to constrain the hash code distance of similar images; Calculate the quantization loss to reduce the error of the binarization process; The weighted negative log-likelihood loss, the metric alignment loss, and the quantization loss are fused to obtain a total loss function.
9. The deep hash image retrieval method based on feature fusion and dynamic transformer according to claim 8 is characterized in that: The step 4 further includes: Taking the total loss function as the optimization target, the model parameters are gradient calculated by the standard back-propagation algorithm and iteratively updated using the Adam optimizer; This process is executed repeatedly until the total loss function converges or the preset number of training rounds is reached, thereby completing the training.
10. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 9.
Citation Information
Cited By
OCT image choroidal neovascularization segmentation method and system
CN120997226A