Similar image retrieval method and system based on deep hashing model of binary network

CN119202289BActive Publication Date: 2026-09-25UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411240704.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-09-25
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

[0004]但是,目前的语义哈希方法仅依赖哈希码的特性在搜索过程中加速,却忽略了模型在推断过程中的计算效率问题

Benefits of technology

[0018]由上述本发明提供的技术方案可以看出,构建了一个基于二值网络的深度哈希模型并用于对图像数据进行快速检索,相比现有方案,本发明考虑了模型推断过程中的计算开销问题,基于二值网络,大幅减少了模型的存储大小以及计算的复杂度,从而能够部署在资源受限的设备上,同时因为二值网络的表达能力受限,本发明提出一种基于语义的哈希中心自适应模块以及一个两阶段的训练方法,确保检索结果的有效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202289B_ABST
    Figure CN119202289B_ABST
Patent Text Reader

Abstract

The application discloses a similar image retrieval method and system based on a deep hash model of a binary network, and provides a scheme in which a deep hash model based on a binary network is constructed and used for fast retrieval of image data; compared with an existing scheme, the scheme provided by the application considers the calculation overhead problem in the model inference process, is based on a binary network, greatly reduces the storage size of the model and the calculation complexity, and can be deployed on a resource-limited device; meanwhile, because the expression capability of the binary network is limited, a hash center adaptive module based on semantics and a two-stage training method are further provided, so that the effectiveness of the retrieval result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of similar image retrieval technology, and in particular to a similar image retrieval method and system based on a deep hash model of binary networks. Background Technology

[0002] In recent years, with the rapid development of online technology, image data has increased exponentially, giving users access to massive amounts of image data. Similar image retrieval systems (or content-based image retrieval systems) can find images similar to the query image from a vast pool of image resources, greatly alleviating information overload and thus being widely used in various applications. Furthermore, with rapid technological advancements, similar image retrieval systems have undergone improvements from manual feature-based methods to deep learning methods, continuously enhancing retrieval accuracy. However, retrieval efficiency remains an unavoidable key issue in similar image retrieval systems, closely related to user experience and system load. Therefore, improving retrieval efficiency is a pressing research problem that needs to be addressed in similar image retrieval systems.

[0003] To address this research problem, researchers have proposed various approaches, among which deep hashing has garnered significant attention in recent years. Its main idea is to use deep network models such as ResNet and VGG, along with binarization operations, to map image data into binary representations, also known as hash codes. Images are then stored in hash buckets corresponding to these hash codes. On one hand, the similarity between two hash codes can be quickly calculated using computer instructions such as XOR and POPCOUNT. On the other hand, semantic hashing indexes can be used to quickly return relevant images from a candidate image set. Therefore, semantic hashing technology can significantly alleviate the efficiency problem of search.

[0004] However, current semantic hashing methods rely solely on the properties of hash codes to accelerate the search process, neglecting the computational efficiency of the model during inference. Specifically, existing semantic hashing works often use large model skeletons to extract image features to ensure search accuracy, which increases the computational overhead of the deep hashing model and makes it difficult to deploy on resource-constrained devices. Therefore, reducing the computational load of the deep hashing model's inference process while maintaining retrieval quality is a crucial issue in deep hashing model design. Summary of the Invention

[0005] The purpose of this invention is to provide a similar image retrieval method and system based on a deep hash model using binary networks. This method can reduce the computational load of the deep hash model inference process while ensuring retrieval quality, and thus can be applied to resource-constrained devices.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] A similar image retrieval method based on a deep hashing model using binary networks includes:

[0008] Construct a deep hashing model based on binary networks;

[0009] A semantic-based hash center adaptive module is introduced to train the binary network-based deep hashing model. This module includes a classification model to obtain semantic information about categories and generate hash centers. The training process consists of two stages: In the first stage, the classification loss of the classification model is used to train the semantic-based hash center adaptive module; the second stage includes a spatial exploration stage and a code aggregation stage. In the spatial exploration stage, the hash center alignment loss is calculated using the hash centers and the binary-like hash representations output by the binary network-based deep hashing model, and combined with the classification loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module. In the code aggregation stage, the binary network-based deep hashing model obtains a set of binary-like hash representations, calculates the centroid vector of each semantic category, and processes it using an extrapolation method. The extrapolation aggregation loss is calculated using the processing results, and combined with the classification loss and the hash center alignment loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module.

[0010] The trained binary network-based deep hash model is used to obtain a set of binary hash representations of image data in the database and to construct a semantic hash index. The trained binary network-based deep hash model is used to obtain the hash code of the query image and to perform image retrieval in combination with the semantic hash index.

[0011] A similar image retrieval system based on a deep hashing model using binary networks, comprising:

[0012] Model building unit, used to build a deep hashing model based on a binary network;

[0013] The model training unit is used to introduce a semantic-based hash center adaptive module to train the binary network-based deep hashing model. This module includes a classification model to obtain semantic information about categories and generate hash centers. The training process consists of two stages: In the first stage, the classification loss of the classification model is used to train the semantic-based hash center adaptive module; the second stage includes a space exploration stage and a code aggregation stage. In the space exploration stage, the hash center alignment loss is calculated using the hash centers and the binary-like hash representations output by the binary network-based deep hashing model, and combined with the classification loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module. In the code aggregation stage, the binary network-based deep hashing model obtains a set of binary-like hash representations, calculates the centroid vector of each semantic category, and processes it using an extrapolation method. The extrapolation aggregation loss is calculated using the processing results, and combined with the classification loss and hash center alignment loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module.

[0014] The image retrieval unit is used to obtain a set of binary hash representations of image data in the database using a trained deep hash model based on a binary network, and to construct a semantic hash index; and to obtain the hash code of the query image using the trained deep hash model based on a binary network, and to perform image retrieval in conjunction with the semantic hash index.

[0015] A processing device includes: one or more processors; and a memory for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0017] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.

[0018] As can be seen from the technical solution provided by the present invention, a deep hashing model based on a binary network is constructed for fast retrieval of image data. Compared with existing solutions, the present invention considers the computational overhead in the model inference process. Based on the binary network, the storage size and computational complexity of the model are greatly reduced, thus enabling deployment on resource-constrained devices. At the same time, because the expressive power of the binary network is limited, the present invention proposes a semantic-based hash center adaptive module and a two-stage training method to ensure the effectiveness of the retrieval results. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a similar image retrieval method based on a deep hashing model using a binary network, provided as an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of a similar image retrieval method based on a deep hash model using a binary network, provided in an embodiment of the present invention.

[0022] Figure 3 A schematic diagram illustrating a training scheme for a deep hashing model based on a binary network, provided in an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of a similar image retrieval system based on a deep hash model using a binary network, provided in an embodiment of the present invention;

[0024] Figure 5 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0026] First, the following explanations are provided for the terms that may be used in this article:

[0027] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".

[0028] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0029] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0030] The following provides a detailed description of a similar image retrieval method and system based on a deep hashing model using binary networks, as provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they should be performed according to conventional conditions in the art or conditions recommended by the manufacturer.

[0031] Example 1

[0032] This invention provides a similar image retrieval method based on a deep hashing model using binary networks, such as... Figure 1 As shown, it mainly includes the following steps:

[0033] Step 1: Construct a deep hashing model based on binary networks.

[0034] In this embodiment of the invention, the deep hashing model based on a binary network includes: an image feature extractor of a binary network and a hash layer; wherein, the image feature extractor extracts features v of the image through the binary network, and the hash layer obtains a binary hash-like representation. During the training phase, a binary-like hash representation is used for training. After training is complete, the binary-like hash representation is converted into a binary hash code h using a symbol function.

[0035] Step 2: Introduce a semantic-based hash center adaptive module to train the deep hash model based on binary networks.

[0036] In this embodiment of the invention, the semantic-based hash center adaptive module includes a classification model. The classification model obtains semantic information of categories and generates hash centers. The training process includes two stages. In the first stage, the classification loss of the classification model is used to train the semantic-based hash center adaptive module. The second stage includes a spatial exploration stage and a code aggregation stage. In the spatial exploration stage, the hash center alignment loss is calculated using the hash center and the quasi-binary hash representation output by the deep hash model based on a binary network. This loss is then combined with the classification loss to train the deep hash model based on a binary network and the semantic-based hash center adaptive module. In the code aggregation stage, the set of quasi-binary hash representations is obtained using the deep hash model based on a binary network. The centroid vector of each semantic category is calculated and processed using an extrapolation method. The extrapolation aggregation loss is calculated using the processing result. This loss is then combined with the classification loss and the hash center alignment loss to train the deep hash model based on a binary network and the semantic-based hash center adaptive module.

[0037] Step 3: Perform image retrieval using the trained deep hash model based on a binary network.

[0038] In this embodiment of the invention, a set of binary hash representations of image data in the database is obtained by using a trained deep hash model based on a binary network, and a semantic hash index is constructed; and the hash code of the query image is obtained by using the trained deep hash model based on a binary network, and image retrieval is performed in conjunction with the semantic hash index.

[0039] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.

[0040] I. Overview of Principles

[0041] This invention provides a similar image retrieval scheme based on a deep hashing model using binary networks, overcoming the problems of high computational cost and inapplicability to resource-constrained devices in the deployment of existing deep hashing models. This invention provides a deep hashing model based on binary networks, primarily used to generate semantically informative hash representations (binary hash-like representations) from image data, reducing computational overhead during the generation process, thereby expanding the application scenarios of semantic hashing, such as its application in resource-constrained devices. Based on current deep hashing models, this invention allows the replacement of the original image representation extraction module with any binary network. However, due to the weak representational power of binary networks, the accuracy of the generated hash codes is significantly reduced compared to the original real-valued networks. Therefore, this invention, based on the characteristics of binary networks, designs a semantically adaptive hash center module and a two-stage training method to ensure the final retrieval effect.

[0042] During training, image data and their corresponding semantic labels are acquired from the internet, and invalid and noisy data are removed. The data is then divided into training and validation sets.

[0043] In the first stage, the semantic-based hash center module is trained, using cross-entropy as the target loss. The semantic center module in this invention can use any image feature extraction network, such as the ResNet series, VGG series, or ViT series networks. In this invention, ResNet50 is used as the image feature extraction network. Through the trained semantic-based hash center module, a hash center can be generated for each semantic meaning.

[0044] The second stage involves training a deep hashing model based on a binary network. This invention generates a hash center from a semantic-based hash center module as the target for the corresponding semantic image, employing a two-stage training approach to leverage the unique properties of binary networks in deep hashing. In the spatial exploration stage, a hash center loss is used to encourage hash codes to approach the hash center. In the code aggregation stage, an additional extrapolation aggregation loss objective is added to encourage the aggregation of hash codes of the same category. The binary network used in this embodiment can be any currently available binary network; ReActNet is used in this embodiment, but other binary networks are also applicable.

[0045] like Figure 2 The diagram illustrates an example of the overall process of this invention: In the offline stage, using the training method proposed in this invention, after training, the image data in the database is processed through a trained hash code generation model based on a binary network to obtain a corresponding hash code set, and this hash code set is used to construct a semantic hash index. In the online stage, the query image is processed through a hash code generation model based on a binary network to obtain a hash code, and the semantic hash index is used to quickly find the relevant image and return the result.

[0046] II. Detailed introduction of the plan.

[0047] 1. Problem definition and formalization.

[0048] In this embodiment of the invention, two definitions need to be formalized: semantic hashing task based on binary networks and semantic hashing retrieval framework.

[0049] Semantic hashing task based on binary networks: Define the set of images as X = {x1, x2, ... x...} N}, where x j Let represent the j-th image, where j = 1, 2, ..., N, and N represents the total number of images. The goal of semantic hashing based on binary networks is to learn a hash function f(x): x → h, which maps image x to a binary representation h ∈ {-1, 1}.b The symbol 'b' represents the length of the binary representation. This binary representation 'h' is also called the hash code. The hash function is implemented using a binary network.

[0050] Semantic hashing retrieval framework: The goal of the semantic hashing retrieval framework is to find a set of images S(q) similar to a user-given query image q. This framework consists of two phases: an offline phase and an online phase. In the offline phase, a hash model f(x): x→h, x∈X based on a binary network is used to map the image set X to hash codes and construct a hash index. In the online phase, the same hash model f(x): x→h is used. q Obtain the hash code of the query and quickly find the relevant image using the hash index.

[0051] 2. Data collection and preprocessing.

[0052] This invention uses images as the input dataset. Such datasets include CIFAR100, IMAGENET, and MSCOCO. Alternatively, various image datasets can be collected via web scraping or offline methods as input data.

[0053] 3. Model building and training.

[0054] like Figure 3 As shown, the structure of the deep hashing model based on binary networks and the semantic-based hash center adaptive module, along with the related training scheme, are illustrated. The semantic-based hash center adaptive module generates hash centers using the category semantics of the image and adaptively adjusts them during training. The first stage of training mainly focuses on training the semantic-based hash center adaptive module. The second stage is divided into a spatial exploration training stage and a hash code aggregation training stage. The spatial exploration training stage uses a hash center loss objective to encourage hash codes to approach their corresponding hash centers. The hash code aggregation training stage adds an extrapolation aggregation loss objective, aiming to cluster hash codes of the same category more tightly in the hash space while maintaining sufficient distance between hash codes of different categories.

[0055] (1) A deep hashing model based on binary networks.

[0056] like Figure 2 As shown in part (a), the deep model based on binary networks consists of an image feature extractor of a binary network and a hash layer.

[0057] For example, a ReactNet model H(x):x→v can be used as a feature extractor, where v represents the image features. The features v are then passed through a hash layer to obtain a hash-like representation.

[0058] For example, a hash layer may consist of a single-layer feedforward network and a tanh(·) activation function.

[0059] The above process can be formalized as follows:

[0060]

[0061] Among them, W h and b h This represents the learnable parameters in the hash layer.

[0062] During the training phase of the model, a binary-like hash representation is used. Participate in training to ensure gradients can propagate in the correct direction. During the inference phase after model deployment, [the following will be done / will be implemented / etc.]. The final binary hash code is obtained through a sign function, `sign`.

[0063] It is worth noting that if traditional deep hashing training methods are used, the accuracy of the hash code will be significantly reduced. This is because the representational power of binary networks is significantly lower than that of real-valued networks. Therefore, this invention introduces a special module and training method, which will be described in detail later.

[0064] (2) Semantic-based hash center adaptive module and its training scheme.

[0065] The main goal of the semantic-based hash center adaptive module is to determine the hash center based on the image x. i Category information y i ∈{0,1} m Obtain a corresponding hash center c′ i ={-1,1} b This better guides the distribution of generated hash codes in space, where m represents the total number of categories in the data. A direct approach is to use the label y i Input a neural network, such as a multilayer perceptron (MLP), to obtain the hash center c′. i However, this method ignores the semantic information in the categories. Therefore, this embodiment of the invention introduces a semantic-based hash center adaptive module, which makes the generated hash center more accurate by incorporating the semantic information of the categories.

[0066] First, a matrix of category semantics is obtained through a pre-training method. Where R is the symbol for the set of real numbers, and C m Represents the category semantic vector w lThe dimensions are l = 1, 2, ..., m. In this embodiment of the invention, the category semantic matrix is ​​obtained through a classification task. The semantic-based hash center adaptive module includes a classification model, such as... Figure 2 In the example shown in part (b), a ResNet50 network can be used as the image classification model. The last layer is the class probability prediction layer, and the second-to-last layer is the semantic layer for the class. Therefore, the network parameters of the second-to-last layer can be regarded as the semantic layer W of the class. Assume the training batch data is... Where |B| represents the size of the training batch. A classification task is set to train the network, enabling the category semantic layer to acquire the semantics of the category:

[0067]

[0068] in, This represents the predicted probability of the classification model for the l-th class; it is the prediction result output by the classification model. The l-th element, y il It is category information y i The l-th element, i.e., the image x i The true information belonging to the l-th category. Where y i It is represented in the form of a bag-of-words model.

[0069] After training the classification model, the parameters of the penultimate layer can be used as the matrix of category semantics. w l Let x be the semantic vector of the l-th category. For the i-th image x in training batch B... i , and its category information y i After performing matrix multiplication with the category semantic matrix W, the corresponding category feature w′ is obtained. i :

[0070] w′ i =y i W

[0071] The category features w′ are processed by a feedforward neural network and an activation function. i Mapped to a probability vector p with the same dimension as the binary hash representation. i The probability vector p i Using parameters from the multidimensional Bernoulli distribution, we sample to obtain the hash center c′. i ={-1,1} b Where b represents the dimension of the binary hash representation. Since the sampling process is not differentiable, a reparameterization method is used, and the complete representation of the relevant process is:

[0072] c′ i=sign(σ(F(w′) i ))-μ)

[0073] Here, F represents a feedforward neural network, which can be implemented using an MLP, for example. σ represents the sigmoid activation function. sign represents the sign function, mapping values ​​greater than or equal to 0 to 1 and values ​​less than 0 to -1. μ represents a value obtained from a uniform distribution U(0,1). In this representation, the derivative of sign is 0, thus preventing the derivative from being propagated to the model. In this embodiment, a straight-forward estimator (STE) is used to approximate the gradient, enabling the module to be trained end-to-end.

[0074] (3) Training plan for the second stage.

[0075] In this embodiment of the invention, the second stage is divided into a spatial exploration stage and a code aggregation stage. These two stages are derived from an experimental observation during the research process of this invention. When using a binary network as the image feature extraction model for deep hashing, hash codes corresponding to images of the same category will first disperse and then aggregate during the training process. Based on this characteristic, this invention divides the training process of the deep hashing model based on a binary network into a spatial exploration training stage and a hash code aggregation training stage, and designs different training objectives for different stages.

[0076] (3.1) Space exploration training phase.

[0077] During the spatial exploration training phase, in order to conform to the training characteristics of binary networks, only the image x is considered. i generated hash code To get closer to the corresponding hash center c′ i In each training batch, image x is first... i Inputting a deep hash model based on a binary network yields the corresponding binary-like hash representation. Then, the category semantic matrix W is used as input, and the hash center representation C = {c1, c2, ..., c3} of each semantic information is obtained through the feedforward neural network F of the semantic-based hash center adaptive module, the sigmoid function, and the sign function. m The specific process is as follows:

[0078] [c1,c2,…,c m ] T =sign(σ(F(W))-0.5)

[0079] Where T is the transpose symbol, and the sigmoid function σ represents the application of hash center alignment to each row of the matrix. Then, the following hash center alignment loss objective is used:

[0080]

[0081] Among them, L hc Let |B| be the hash center alignment loss, |B| be the number of images in training batch B, and φ be the similarity calculation function (e.g., cosine similarity). For the i-th image x in the training batch i The binary hash representation of c′ i For the i-th image x i The hash center, τ represents the temperature hyperparameter, c l ∈C represents the hash center representation of l categories.

[0082] (3.2) Code aggregation stage.

[0083] During the code aggregation phase, in addition to using the hash center mentioned earlier to align the loss target L... hc To encourage hash codes to approach their hash centers, an additional extrapolation aggregation loss objective is introduced to enhance aggregation between hash codes corresponding to images of the same semantic category. A basic idea is to train a batch of data... The hash code distance between images of the same category is reduced. Here, |B| represents the size of the training batch. Another consideration is ensuring that samples of the same category within a batch are mutually difficult samples, thereby further promoting the aggregation of hash codes. To this end, this invention proposes an extrapolation aggregation loss objective.

[0084] Specifically: regarding data Several sets of augmented views are obtained through data processing, and all sets of augmented views are used as the training set B. new The training set B is obtained through a deep hashing model based on a binary network. new The binary hash representation set H new ,definition H represents new The set of binary hash representations belonging to class l is used to compute the binary hash representation set H. new Centroid vector of each semantic category m represents the total number of categories.

[0085] like Figure 3 As shown in section (c), data augmentation is used for each image x. i =B generates two views and In this example, B new =B1∪B2. Of course, in practical applications, this can be extended to more views, and this invention does not limit it.

[0086] Wherein, the centroid vector s corresponding to the l-th categoryl The calculation method is as follows:

[0087]

[0088] in, express The number of elements in for One of the elements in the array is a binary hash representation;

[0089] The extrapolation method is used to process the data to obtain the desired results. The new representation of h′ k The Hard Positive Sample Representation is represented as follows:

[0090]

[0091] Where Φ represents the extrapolation method, s l for The centroid vector corresponding to the semantic category, α ~ U(0,1)+1, means that α is the value sampled from U(0,1)+1, where U(0,1) represents a uniform distribution with an expectation of 0 and a variance of 1. Adding 1 at the end ensures that the range of α is between [1,2]. Through the above settings, we can ensure... Thus increase The variance of the distance between two binary hashes is represented by a binary hash.

[0092] Finally, the extrapolated aggregation loss is calculated using the processing results and expressed as:

[0093]

[0094] Among them, L asd To extrapolate the aggregation loss, for Another element in; D h This is the distance calculation function.

[0095] The above extrapolation aggregation loss not only effectively utilizes information from other images sharing the same category within a batch, but also increases the variance of the hash code, thereby improving the performance of hash code aggregation. Furthermore, when the batch size |B| is less than the number of categories m, some categories may only have one image or a small number of images in each training batch. To address this issue, this invention employs a stratified sampling approach to construct training batches. That is, first, some categories are sampled, then images are sampled within these categories, and finally, a training batch is constructed.

[0096] (4) Training process of the complete model.

[0097] The first stage involves applying the objective function L to the classification model in the semantic-based hash center adaptive module. cls Training is performed to ensure that the category semantic matrix W has basic category semantic information.

[0098] In the second phase:

[0099] During the spatial exploration training phase, the semantic-based hash center adaptive module and the binary network-based deep hash model are trained together using the following loss function:

[0100] L2 = L hc +β1L cls

[0101] Set the hyperparameter T, and after T training steps, enter the code aggregation training phase, adding the extrapolation aggregation loss target L. asd The following loss function is obtained:

[0102] L3 = L hc +β1L cls +β2L asd

[0103] Among them, β1 and β2 are set parameters used to control the proportion of related losses.

[0104] Since the specific process of model training based on loss function can be implemented by referring to conventional techniques, it will not be elaborated here.

[0105] 4. Semantic hash index construction.

[0106] After training, the deep hashing model based on a binary network is deployed to a suitable device. The candidate image set X = {x1, x2, ... x...} is then used. N The image is scaled to a predefined size and then input into a trained binary network-based deep hashing model to obtain the hash code of the candidate image set. Then, a semantic hash index ψ(x) is constructed using the Multi-Index Hashing method.

[0107] 5. Online search.

[0108] When used online, the system receives a query image q from the user and first scales the image to a predefined size. Then, it feeds the image into a trained deep hashing model based on a binary network to obtain the corresponding hash code h. q The hash code h of the query image q is obtained using a trained deep hashing model based on a binary network. q Set a search threshold K and perform image retrieval as follows:

[0109] Step (1): Initialize the search radius r = 0, and the search result is R;

[0110] Step (2): Look up the hash code h using the semantic hash index. q For images with a distance of r, the retrieved images are placed into the search results R;

[0111] Step (3): Calculate whether the number of images in set R is less than the retrieval threshold K. If yes, let r = r + 1 and jump to step (2); if no, go to step (4).

[0112] Step (4): Obtain the search result R, which contains the similar images that were retrieved.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0114] Example 2

[0115] This invention also provides a similar image retrieval system based on a deep hashing model using binary networks, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 4 As shown, the system mainly includes:

[0116] Model building unit, used to build a deep hashing model based on a binary network;

[0117] The model training unit is used to introduce a semantic-based hash center adaptive module to train the binary network-based deep hashing model. This module includes a classification model to obtain semantic information about categories and generate hash centers. The training process consists of two stages: In the first stage, the classification loss of the classification model is used to train the semantic-based hash center adaptive module; the second stage includes a space exploration stage and a code aggregation stage. In the space exploration stage, the hash center alignment loss is calculated using the hash centers and the binary-like hash representations output by the binary network-based deep hashing model, and combined with the classification loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module. In the code aggregation stage, the binary network-based deep hashing model obtains a set of binary-like hash representations, calculates the centroid vector of each semantic category, and processes it using an extrapolation method. The extrapolation aggregation loss is calculated using the processing results, and combined with the classification loss and hash center alignment loss to train both the binary network-based deep hashing model and the semantic-based hash center adaptive module.

[0118] The image retrieval unit is used to obtain a set of binary hash representations of image data in the database using a trained deep hash model based on a binary network, and to construct a semantic hash index; and to obtain the hash code of the query image using the trained deep hash model based on a binary network, and to perform image retrieval in conjunction with the semantic hash index.

[0119] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0120] Example 3

[0121] The present invention also provides a processing device, such as Figure 5 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0122] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0123] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example:

[0124] Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc.

[0125] The output device can be a display terminal;

[0126] The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.

[0127] Example 4

[0128] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.

[0129] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0130] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A similar image retrieval method based on a deep hashing model using binary networks, characterized in that, include: Construct a deep hashing model based on binary networks; A semantic-based hash center adaptive module is introduced to train the binary network-based deep hash model. This module includes a classification model to obtain semantic information about categories and generate hash centers. Specifically, the penultimate layer of the classification model is a semantic layer for the categories, and the network parameters of this penultimate layer are used as the semantic information of the categories, i.e., the category semantic matrix. , For the first The semantic vectors of each category, , This represents the total number of categories; for the i-th image... Its category information , and the matrix of category semantics The corresponding categorical features are obtained after performing matrix multiplication. : ;Class features are combined using a feedforward neural network and activation function Mapped to a probability vector of the same dimension as a binary hash representation. , this probability vector The parameters of the multidimensional Bernoulli distribution are used to sample the data to obtain the hash center. ,in, The dimension of the binary hash representation is represented. The training process consists of two stages. In the first stage, the classification loss of the classification model is used to train the semantic hash center adaptive module. The second stage includes a space exploration stage and a code aggregation stage. In the space exploration stage, the hash center alignment loss is calculated using the hash center and the binary hash representation output by the deep hash model based on the binary network. The classification loss is then combined with the deep hash model based on the binary network and the semantic hash center adaptive module to train the module. In the code aggregation stage, the set of binary hash representations is obtained using the deep hash model based on the binary network. The centroid vector of each semantic category is calculated and processed using an extrapolation method. The extrapolation aggregation loss is calculated using the processing result. The classification loss and the hash center alignment loss are then combined with the hash center alignment loss to train the deep hash model based on the binary network and the semantic hash center adaptive module. The step of obtaining a set of binary hash representations using a deep hash model based on binary networks, calculating the centroid vector of each semantic category, and processing it using an extrapolation method includes: Record the data of the current training batch. Several sets of augmented views are obtained through data processing, and all sets of augmented views are used as the training set. The training set is obtained through a deep hashing model based on binary networks. binary hash representation set ,definition express Belongs to the category The set of binary hash representations is calculated. Centroid vector of each semantic category , Indicates the total number of categories; Among them, the Centroid vectors corresponding to each category The calculation method is as follows: ; in, express The number of elements in for One of the elements in the array is a binary hash representation; The extrapolation method is used to process the data to obtain the desired results. new representation , is represented as: ; in, Indicates the extrapolation method. for The centroid vector corresponding to the semantic category, ,Right now From The value obtained from sampling This represents a uniform distribution with an expected value of 0 and a variance of 1. The trained binary network-based deep hash model is used to obtain a set of binary hash representations of image data in the database and to construct a semantic hash index. The trained binary network-based deep hash model is used to obtain the hash code of the query image and to perform image retrieval in combination with the semantic hash index.

2. The similar image retrieval method based on a deep hashing model using a binary network according to claim 1, characterized in that, The deep hashing model based on binary networks includes: an image feature extractor of a binary network and a hash layer; wherein, the image feature extractor extracts image features through the image feature extractor of the binary network. And obtain a binary hash representation through a hash layer. During the training phase, a binary-like hash representation is used for training. After training, the binary-like hash representation is converted into a binary hash code using a sign function. .

3. The similar image retrieval method based on a deep hashing model using a binary network according to claim 1, characterized in that, The calculation of hash center alignment loss using the hash center and the binary-like hash representation output by the deep hash model based on binary networks includes: The matrix of category semantics As input, a hash center representation of each semantic information is obtained through a feedforward neural network F, an activation function, and a sign function of a semantically based hash center adaptive module. The process is as follows: ; Where T is the transpose sign, and the activation function is... This indicates that it is used for each row of the matrix. Indicates the first The hash center representation corresponding to each category; The hash center alignment loss is then calculated using the following formula: ; in, For hash center alignment loss, For the current training batch The number of images in For similarity calculation function, For the i-th image The binary hash representation, For the i-th image hash center This indicates the temperature hyperparameter.

4. The similar image retrieval method based on a deep hashing model using a binary network according to claim 1, characterized in that, The extrapolated aggregation loss, calculated using the processing results, is expressed as follows: ; in, To extrapolate the aggregation loss, for Another element in; This is the distance calculation function.

5. The similar image retrieval method based on a deep hashing model using a binary network according to claim 1, characterized in that, The step of obtaining the hash code of the query image using a trained binary network-based deep hashing model and performing image retrieval in conjunction with a semantic hash index includes: The query image is obtained using a trained deep hashing model based on a binary network. hash code Set a search threshold K and perform image retrieval as follows: Step (1): Initialize the search radius r=0, and the search result is R; Step (2): Look up the hash code using the semantic hash index. For images with a distance of r, the retrieved images are placed into the search results R; Step (3): Calculate whether the number of images in set R is less than the retrieval threshold K. If yes, let r = r + 1 and jump to step (2); if no, jump to step (4). Step (4): Obtain the search result R, which contains the similar images that were retrieved.

6. A similar image retrieval system based on a deep hashing model using binary networks, characterized in that, include: Model building unit, used to build a deep hashing model based on a binary network; The model training unit is used to train the deep hash model based on a binary network by introducing a semantic-based hash center adaptive module. The semantic-based hash center adaptive module includes a classification model, which obtains semantic information of categories and generates hash centers. This includes: the penultimate layer of the classification model is a semantic layer for the categories, and the network parameters of the penultimate layer are used as the semantic information of the categories, i.e., the category semantic matrix. , For the first The semantic vectors of each category, , This represents the total number of categories; for the i-th image... Its category information , and the matrix of category semantics The corresponding categorical features are obtained after performing matrix multiplication. : ;Class features are combined using a feedforward neural network and activation function Mapped to a probability vector of the same dimension as a binary hash representation. , this probability vector The parameters of the multidimensional Bernoulli distribution are used to sample the data to obtain the hash center. ,in, The dimension of the binary hash representation is represented. The training process consists of two stages. In the first stage, the classification loss of the classification model is used to train the semantic hash center adaptive module. The second stage includes a space exploration stage and a code aggregation stage. In the space exploration stage, the hash center alignment loss is calculated using the hash center and the binary hash representation output by the deep hash model based on the binary network. The classification loss is then combined with the deep hash model based on the binary network and the semantic hash center adaptive module to train the module. In the code aggregation stage, the set of binary hash representations is obtained using the deep hash model based on the binary network. The centroid vector of each semantic category is calculated and processed using an extrapolation method. The extrapolation aggregation loss is calculated using the processing result. The classification loss and the hash center alignment loss are then combined with the hash center alignment loss to train the deep hash model based on the binary network and the semantic hash center adaptive module. The process of obtaining a set of binary hash representations using a deep hash model based on binary networks, calculating the centroid vector of each semantic category, and processing it using an extrapolation method includes: Record the data of the current training batch. Several sets of augmented views are obtained through data processing, and all sets of augmented views are used as the training set. The training set is obtained through a deep hashing model based on binary networks. binary hash representation set ,definition express Belongs to the category The set of binary hash representations is calculated. Centroid vector of each semantic category , Indicates the total number of categories; Among them, the Centroid vectors corresponding to each category The calculation method is as follows: ; in, express The number of elements in for One of the elements in the array is a binary hash representation; The extrapolation method is used to process the data to obtain the desired results. new representation , is represented as: ; in, Indicates the extrapolation method. for The centroid vector corresponding to the semantic category, ,Right now From The value obtained from sampling This represents a uniform distribution with an expected value of 0 and a variance of 1. The image retrieval unit is used to obtain a set of binary hash representations of image data in the database using a trained deep hash model based on a binary network, and to construct a semantic hash index; and to obtain the hash code of the query image using the trained deep hash model based on a binary network, and to perform image retrieval in conjunction with the semantic hash index.

7. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 5.

8. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.