Image retrieval method and device based on similarity, storage medium and electronic device
Through model training based on the similarity between text and image features, the problems of high cost and low efficiency of image annotation in the prior art are solved, and an efficient image retrieval method is realized, which is suitable for smart homes and smart home scenes.
Patent Information
- Application Number
- CN202311855151.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing text-based image retrieval methods require a lot of manpower and time to mark images, and the labeling quality and consistency are difficult to guarantee, and the text information cannot fully cover the image content, resulting in intricate and cross-domain image retrieval efficiency.
Through model training based on the similarity between text and image features, the text feature extraction network and image feature extraction network are used to obtain the target feature vectors of the search information, and the candidate feature vectors with similarity satisfies the preset threshold from the vector library, so as to realize searching pictures with text and searching pictures with pictures.
It reduces the cost of manual labeling, improves the efficiency and accuracy of image retrieval, and is suitable for complex and cross-domain image retrieval tasks.
Smart Images

Figure CN120234431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of smart homes, and particularly to an image retrieval method, device, storage medium, and electronic device based on similarity. Background Art
[0002] The text-based image retrieval method refers to a technology that uses text information to describe and query images. It usually requires manual or semi-automatic annotation of images, extracting metadata such as the name, size, type, author, and age of the image, or content such as objects, scenes, and colors in the image. Then, according to the keywords or classification catalogs input by the user, images that match the text information are retrieved from the image database. The advantages of the text-based image retrieval method are simple to use and can utilize existing text retrieval technologies without analyzing the visual elements of the images.
[0003] However, it also has some disadvantages. For example, it requires a large amount of human and time costs to annotate images, and it is difficult to ensure the quality and consistency of the annotation; text information often cannot fully cover all the content of the image, resulting in some important information being ignored or blurred; there is a semantic gap between text information and image information, and the same word may correspond to multiple different images, or the same image may correspond to multiple different words.
[0004] Based on this, the text-based image retrieval method is suitable for some simple, metadata-based, or specific-domain image retrieval tasks, but for some complex, content-based, or cross-domain image retrieval tasks, its effect and efficiency are limited. Summary of the Invention
[0005] The purpose of this application is to provide an image retrieval method, device, storage medium, and electronic device based on similarity, which are used to train a model through the feature similarity between images and text, so that the model can perform image search by text and image search by image based on the feature similarity.
[0006] This application provides an image retrieval method based on similarity, including:
[0007] Obtain retrieval information; the retrieval information includes: retrieval text and retrieval image; input the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; screen out multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets a preset similarity, and use the image corresponding to each candidate feature vector among the multiple candidate feature vectors as the retrieval result image corresponding to the retrieval information.
[0008] Optionally, the step of inputting the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information includes: inputting the retrieval text into the text feature extraction model for feature extraction to obtain a first feature vector; inputting the first feature vector into the feedforward neural network for feature dimension adjustment to obtain a second feature vector, and determining the second feature vector as the target feature vector; wherein, the dimension of the second feature vector is the same as the dimension of the image feature vector output by the image feature extraction network; the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network.
[0009] Optionally, the step of inputting the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information includes: inputting the retrieval image into the multiple feature extraction sub-networks for sequential feature extraction, and respectively inputting the features output by each feature extraction sub-network among the multiple feature extraction sub-networks into the residual block for processing to obtain the image vector feature of the retrieval image; determining the image vector feature of the retrieval image as the target feature vector; wherein, the image feature extraction network includes: multiple serially connected feature extraction sub-networks, and a residual block; each feature extraction sub-network among the multiple serially connected feature extraction sub-networks includes: multiple convolutional layers matching the number of color channels of the input image, and a max pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network among the multiple serially connected feature extraction sub-networks are all input into the residual block for processing to obtain the image vector feature of the input image.
[0010] Optionally, inputting the retrieved image into the multiple feature extraction sub-networks for feature extraction in sequence includes: respectively inputting multiple third feature vectors output by a convolutional layer in a target feature extraction sub-network into the max pooling layer and the average pooling layer to obtain a max pooling feature vector and an average pooling feature vector corresponding to the multiple third feature vectors; inputting the max pooling feature vector and the average pooling feature vector into an activation layer and a dropout layer, and outputting an image feature vector corresponding to the target feature extraction sub-network; where the target feature extraction sub-network is any one of the multiple target feature extraction sub-networks.
[0011] Optionally, screening out multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets a preset similarity includes: calculating a similarity value between the target feature vector and each to-be-matched feature vector in the vector library, and determining the to-be-matched feature vectors with similarity values greater than a preset similarity threshold as the multiple candidate feature vectors; where the similarity value between the target feature vector and each to-be-matched feature vector in the vector library is calculated based on cosine similarity.
[0012] Optionally, before inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information, the method further includes: constructing multiple positive samples each containing an image and a description text corresponding to the image, and randomly sampling the images and description texts of the multiple positive samples to obtain multiple negative samples each containing an image and a description text not corresponding to the image.
[0013] Optionally, after constructing multiple positive samples each containing an image and a description text corresponding to the image, and randomly sampling the images and description texts of the multiple positive samples to obtain multiple negative samples each containing an image and a description text not corresponding to the image, the method further includes: inputting the image in a target sample into an image feature extraction network of the target model, and inputting the description text in the target sample into a text feature extraction network of the target model to train the target model; adjusting model parameters of the target model based on a loss value of a cosine loss function.
[0014] This application also provides an image retrieval device based on similarity, including:
[0015] An acquisition module for acquiring retrieval information; the retrieval information includes: retrieval text and retrieval images; a feature extraction module for inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and images; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval images; a feature matching module for screening out a plurality of candidate feature vectors from a vector library whose similarity with the target feature vector meets a preset similarity, and using each image corresponding to each candidate feature vector among the plurality of candidate feature vectors as a retrieval result image corresponding to the retrieval information.
[0016] Optionally, the feature extraction module is specifically configured to input the retrieval text into a text feature extraction model for feature extraction to obtain a first feature vector; the feature extraction module is further specifically configured to input the first feature vector into a feedforward neural network for adjusting the feature dimension to obtain a second feature vector, and determine the second feature vector as the target feature vector; wherein, the dimension of the second feature vector is the same as the dimension of the image feature vector output by the image feature extraction network; the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network.
[0017] Optionally, the feature extraction module is specifically configured to input the retrieval image into the plurality of feature extraction subnetworks for sequential feature extraction, and input the features output by each feature extraction subnetwork among the plurality of feature extraction subnetworks into a residual block for processing to obtain an image vector feature of the retrieval image; the feature extraction module is further specifically configured to determine the image vector feature of the retrieval image as the target feature vector; wherein, the image feature extraction network includes: a plurality of serially connected feature extraction subnetworks, and a residual block; each feature extraction subnetwork in the plurality of serially connected feature extraction subnetworks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a max pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction subnetwork in the plurality of serially connected feature extraction subnetworks are input into the residual block for processing to obtain an image vector feature of the input image.
[0018] Optionally, the feature extraction module is specifically configured to input multiple third feature vectors output by the convolutional layer in the target feature extraction sub-network into the max pooling layer and the average pooling layer respectively, to obtain a max pooling feature vector and an average pooling feature vector corresponding to the multiple third feature vectors; the feature extraction module is further specifically configured to input the max pooling feature vector and the average pooling feature vector into an activation layer and a dropout layer, and output an image feature vector corresponding to the target feature extraction sub-network; wherein, the target feature extraction sub-network is any one of the multiple target feature extraction sub-networks.
[0019] Optionally, the feature matching module is specifically configured to calculate similarity values between the target feature vector and each of the to-be-matched feature vectors in the vector library, and determine the to-be-matched feature vectors with similarity values greater than a preset similarity threshold as the multiple candidate feature vectors; wherein, the similarity values between the target feature vector and each of the to-be-matched feature vectors in the vector library are calculated based on cosine similarity.
[0020] Optionally, the apparatus further includes: a sample construction module; the sample construction module is configured to construct multiple positive samples in which the sample includes an image and a description text corresponding to the image, and perform random sampling on the images and description texts of the multiple positive samples to obtain multiple negative samples in which the sample includes an image and a description text not corresponding to the image.
[0021] Optionally, the apparatus further includes: a model training module; the model training module is configured to input the image in the target sample into the image feature extraction network of the target model, and input the description text in the target sample into the text feature extraction network of the target model to train the target model; the model training module is further configured to adjust the model parameters of the target model based on the loss value of the cosine loss function.
[0022] The present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute steps to implement the similarity-based image retrieval method as described in any one of the above through the computer program.
[0023] The present application further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, and when the program runs, it executes steps to implement the similarity-based image retrieval method as described in any one of the above.
[0024] The present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it executes steps to implement the similarity-based image retrieval method as described in any one of the above.
[0025] The image retrieval method, device, storage medium, and electronic device based on similarity provided by this application first obtain retrieval information; the retrieval information includes: retrieval text and retrieval image; then, input the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; finally, screen out multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets the preset similarity, and use the image corresponding to each candidate feature vector among the multiple candidate feature vectors as the retrieval result image corresponding to the retrieval information; in this way, the model is trained through the feature similarity between images and text, enabling the model to perform image search by text and image search by image based on feature similarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with this application, and are used together with the description to explain the principles of this application.
[0027] To more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 is a schematic diagram of the hardware environment of an interaction method for an intelligent device according to an embodiment of this application;
[0029] Figure 2 is a schematic flowchart of the image retrieval method based on similarity provided by this application;
[0030] Figure 3 is one of the schematic diagrams of the model structure provided by this application;
[0031] Figure 4 is another schematic diagram of the model structure provided by this application;
[0032] Figure 5 is a schematic diagram of the structure of the image retrieval device based on similarity provided by this application;
[0033] Figure 6 is a schematic diagram of the structure of the electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] According to one aspect of an embodiment of the present application, a similarity-based image retrieval method is provided. The similarity-based image retrieval method is widely used in smart home (SmartHome), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned similarity-based image retrieval method can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0037] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart drying rack, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.
[0038] The following describes the professional terms involved in the embodiments of the present application:
[0039] Bert model: It is a pre-trained deep bidirectional language representation model. It uses a large amount of unlabeled text and learns the semantic and structural information of the text through the Masked Language Model (MLM) and the Next Sentence Prediction (NSP). Then, the Bert model can be applied to various natural language processing tasks, such as question answering, text classification, named entity recognition, etc., through fine-tuning, and has achieved good results. The main features of the Bert model are: it uses the Transformer model as the basic architecture and uses the self-attention mechanism to capture long-distance dependencies in the text; it is a bidirectional model that can utilize context information on both the left and right sides simultaneously, unlike traditional language models that can only predict from left to right or from right to left; it uses the method of Masked Language Model, that is, randomly covering some words in the input text and then letting the model predict the covered words, so that the model can better focus on the semantics of the entire sentence; it uses the method of Next Sentence Prediction, that is, given two sentences, let the model judge whether they are consecutive, so that the model can better understand the logical relationship between sentences.
[0040] Activation and dropout are two commonly used neural network layers that can improve the performance and generalization ability of the model. The activation layer refers to applying a non-linear function, such as ReLU, sigmoid, or tanh, to the output of neurons to increase the expressive and learning capabilities of the model. The activation layer is usually placed after the convolutional layer or the fully connected layer. The dropout layer refers to randomly discarding the outputs of some neurons during training to reduce overfitting and dependence of the model. The dropout layer is usually placed after the activation layer.
[0041] N-Gram model: It is an algorithm based on statistical language models used to calculate the probability of a piece of text or a sentence. Its basic idea is to perform a sliding window operation of size N on the content in the text by bytes, forming a sequence of byte segments of length N. Each byte segment is called a gram. The occurrence frequencies of all grams are statistically counted and filtered according to a pre-set threshold to form a list of key grams, which is the vector feature space of this text. Each gram in the list is a feature vector dimension.
[0042] Term Frequency-Inverse Document Frequency (TF-IDF): It is a technique used to evaluate the importance of a word in a document collection or a corpus for a certain document. It consists of two parts, term frequency (TF) and inverse document frequency (IDF). Term frequency (TF) refers to the number of times a given word appears in the document, usually normalized to prevent bias towards longer documents. Inverse document frequency (IDF) refers to the reciprocal of the frequency of a word in the corpus, usually taking the logarithm to reduce the influence of words with too high or too low frequencies. The inverse document frequency reflects the discrimination ability of a word. If a word appears in many documents, its discrimination ability is low and the inverse document frequency is small; conversely, if a word appears in only a few documents, its discrimination ability is high and the inverse document frequency is large. The term frequency-inverse document frequency index (TF-IDF) is the product of term frequency (TF) and inverse document frequency (IDF), which represents the importance of a word in a document, and the higher it is, the more important it is. TF-IDF can be used for information retrieval and text mining. By calculating the TF-IDF values of words in different documents, the keywords that best represent the content of the document can be found, or the similarity of different documents can be compared.
[0043] Feed-Forward Neural Network: It is the simplest neural network, consisting of multiple layers of neurons. Each neuron is only connected to the neurons in the previous layer, without feedback or cyclic connections. A feed-forward neural network can accept a set of input signals, and through a series of non-linear transformations, output a set of result signals. Feed-forward neural networks can be used for various machine learning tasks, such as classification, regression, clustering, dimensionality reduction, dimensionality increase, etc.
[0044] Bag-of-Words model (BOW): It is a method for converting text into a vector representation. It ignores the grammar and order of the text and only considers the occurrence times of each word in the text. The basic steps of the bag-of-words model are as follows: Segment the text into several words. Build a dictionary by combining all the non-repeating words that appear in the text into a dictionary. Count the word frequencies, calculate the occurrence times of each word in each text, or use methods such as TF-IDF for weighting. Generate a vector. According to the dictionary and word frequencies, represent each text as a vector with the same length as the dictionary. The bag-of-words model can be used for information retrieval and text mining. By calculating the similarity between the vectors of different texts, tasks such as text classification, clustering, and retrieval can be achieved.
[0045] In view of the above technical problems existing in the related art, the embodiments of the present application provide a similarity-based image retrieval method, which can convert discrete images into continuous vector representations. During the search, the search input is converted into a query vector, and the large-scale vector retrieval technology is used to find the K most matching images, so as to improve the image retrieval efficiency and reduce the labor cost generated by manual annotation.
[0046] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, elaborate in detail on the similarity-based image retrieval method provided by the embodiments of the present application.
[0047] As Figure 2 shown, a similarity-based image retrieval method provided by the embodiments of the present application may include the following steps 201 to step 203:
[0048] Step 201, obtain retrieval information.
[0049] Among them, the retrieval information includes: retrieval text and retrieval image.
[0050] Exemplarily, the above retrieval information is the retrieval information input by the user, and this retrieval information is used to retrieve the images required by the user. For example, taking the retrieval information as text information, the retrieval information can be: high-rise buildings, that is, the user wants to search for images related to high-rise buildings; taking the retrieval information as image information, the retrieval information can be: an image containing high-rise buildings, that is, the user wants to search for images related to high-rise buildings.
[0051] Step 202: Input the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information.
[0052] Among them, the target model is trained based on the feature similarity between text and images; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image. The text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a maximum pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image.
[0053] Exemplarily, as Figure 3 shown, the target model provided in the embodiment of the present application includes: a text feature extraction network and an image feature extraction network, which are respectively used to extract text features and image features. After that, the extracted feature vectors are matched with the feature vectors in the vector library, and the K most similar vectors are selected, and then the corresponding K images, that is, the retrieval result images, are obtained.
[0054] Exemplarily, the above text feature extraction network is mainly composed of a text feature extraction model to extract features from the input text. Since the dimension of the feature vector extracted from the image is different from the dimension of the feature vector extracted from the text, a feedforward neural network FeedForWard is also required in the text feature extraction network to adjust the dimension of the text feature vector extracted by the text feature extraction model in the text feature extraction network.
[0055] In a possible implementation, the above-mentioned text feature extraction network may further include: a bag-of-words model (BOW), a term frequency-inverse document frequency (Tf-idf) algorithm model, and an N-Gram model. Based on the structure of the text feature extraction network, the feature vector corresponding to the term frequency-inverse document frequency and the feature vector corresponding to the N-Gram model can be extracted from the input retrieval text. Then, the two feature vectors are concatenated to obtain a text feature vector. Similar to the text feature extraction network constructed by the text feature extraction model, in the text feature extraction network constructed by the bag-of-words model, the term frequency-inverse document frequency (Tf-idf) algorithm model, and the N-Gram model, a feed-forward neural network (FeedForWard) is also required to adjust the dimension of the extracted text feature vector.
[0056] Exemplarily, based on Figure 3 , such as Figure 4 shown, since an image has three RGB channels, three convolutions can be used to extract features respectively. Then, the extracted features are added to MaxPool (maximum pooling) and AveragePool (average pooling) for further feature extraction. At the same time, activation and dropout neural network layers are also required to improve the performance and generalization ability of the model. This process will go through 3 to 4 times ( Figure 4 in Figure 4 it is 2 times of extraction) to optimize the feature extraction effect. Due to the characteristics of deep neural networks, an overly deep network will cause information loss, so a residual structure is added, that is,
[0057] the residual block in
[0058] Exemplarily, since in the embodiments of the present application, images can be retrieved not only through text but also through images, different processing needs to be performed according to different input types.
[0059] Specifically, taking the retrieval information as the retrieval text as an example, step 202 may include the following steps 202a1 and 202a2:
[0060] Step 202a1: Input the retrieval text into the text feature extraction model for feature extraction to obtain a first feature vector.
[0061] Exemplarily, the text feature extraction model in the embodiments of the present application may be a Bert model.
[0062] Among them, the dimension of the second feature vector is the same as that of the image feature vector output by the image feature extraction network; the text feature extraction network includes: a text feature extraction model and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as that of the image feature vector output by the image feature extraction network.
[0063] Specifically, taking the retrieved information as the retrieved image as an example, step 202 may further include the following steps 202b1 and 202b2:
[0064] Step 202b1: Input the retrieved image into the multiple feature extraction sub-networks, perform feature extraction in sequence, and input the features output by each feature extraction sub-network in the multiple feature extraction sub-networks into the residual block for processing to obtain the image vector feature of the retrieved image.
[0065] Step 202b2: Determine the image vector feature of the retrieved image as the target feature vector.
[0066] Among them, the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks and a residual block; each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a maximum pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks are all input into the residual block for processing to obtain the image vector feature of the input image.
[0067] Specifically, step 202b1 may further include the following steps 202b11 and 202b12:
[0068] Step 202b11: Input the multiple third feature vectors output by the convolutional layer in the target feature extraction sub-network into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature vector and the average pooling feature vector corresponding to the multiple third feature vectors.
[0069] Step 202b12: Input the maximum pooling feature vector and the average pooling feature vector into the activation layer and the dropout layer, and output the image feature vector corresponding to the target feature extraction sub-network.
[0070] Among them, the target feature extraction sub-network is any one of the multiple target feature extraction sub-networks.
[0071] Exemplarily, the operating principles of each feature extraction network in the above steps can be referred to Figure 3 and Figure 4Function descriptions for each part of the model.
[0072] It should be noted that since the sentence lengths of each retrieval text are not uniform, it is necessary to expand the sentences to make all sentences the same length.
[0073] Step 203: Screen out multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets the preset similarity, and use the image corresponding to each candidate feature vector among the multiple candidate feature vectors as the retrieval result image corresponding to the retrieval information.
[0074] Exemplarily, after obtaining the target feature vector corresponding to the retrieval information, image retrieval can be performed based on this target feature vector, and it mainly determines whether it is the image required by the user through the similarity of the feature vectors.
[0075] Specifically, the above step 203 may further include the following step 203a:
[0076] Step 203a: Calculate the similarity values between the target feature vector and each to-be-matched feature vector in the vector library, and determine the to-be-matched feature vectors with similarity values greater than the preset similarity threshold as the multiple candidate feature vectors.
[0077] Among them, the similarity value between the target feature vector and each to-be-matched feature vector in the vector library is calculated based on cosine similarity.
[0078] Exemplarily, in the embodiments of the present application, similarity operations are performed on the obtained target feature vector and each vector in the vector library, and the similarity operations are represented using the Euclidean distance. The formula is:
[0079]
[0080] Among them, n is the number of dimensions, x i and y i are the coordinate values on the i-th dimension respectively.
[0081] Exemplarily, after screening out multiple candidate feature vectors from the vector library whose similarity to the target feature vector is greater than the preset similarity threshold, the image corresponding to each candidate feature vector can be further used as the retrieval result image.
[0082] Optionally, in the embodiments of the present application, the model can be trained with positive and negative samples.
[0083] Exemplarily, before the above step 202, the similarity-based image retrieval method provided by the embodiments of the present application may further include the following steps 204 to 206:
[0084] Step 204: Construct multiple positive samples in the sample that include images and description texts corresponding to the images, and randomly sample the images and description texts of the multiple positive samples to obtain multiple negative samples in the sample that include images and description texts not corresponding to the images.
[0085] Exemplarily, in the embodiments of the present application, it can be defined that a text description is related to a picture as a positive sample, and a text description is not related to a picture as a negative sample. For a picture, we only have positive samples temporarily. We only need to randomly sample from the descriptions of other pictures to obtain negative samples. When randomly sampling negative samples, it is necessary to ensure that the number of negative samples is equal to the number of positive samples, indicating that the number of positive and negative samples is balanced.
[0086] Exemplarily, before training the model using positive and negative samples, it is also necessary to preprocess the description texts in the samples. Data preprocessing mainly includes: full-width to half-width conversion, uppercase numbers to lowercase numbers, uppercase letters to lowercase letters, emoji removal, word segmentation, stop word filtering, etc. Word segmentation is the process of recombining a continuous sequence of characters into a sequence of words according to certain specifications. In English text, spaces are used as natural delimiters between words. We can directly utilize this to perform word segmentation. Since English has tense changes, words after tense transformation need to be transformed back to their original forms.
[0087] Step 205: Input the images in the target samples into the image feature extraction network of the target model, and input the description texts in the target samples into the text feature extraction network of the target model to train the target model.
[0088] Exemplarily, during the training process, it is necessary to perform a similarity operation on the feature vectors obtained from the images and description texts in the samples. The similarity operation uses cosine similarity, and the similarity between them is evaluated by calculating the cosine value of the angle between the two vectors. The formula is:
[0089]
[0090] where A represents the target feature vector, B represents the feature vector to be matched, and i is the dimension of the target feature vector and the feature vector to be matched. A and B have the same dimension.
[0091] Step 206: Adjust the model parameters of the target model based on the loss value of the cosine loss function.
[0092] Exemplarily, in the embodiments of the present application, the loss function of the model is the cosine loss function, and the calculation formula is:
[0093]
[0094] Among them, x1 represents text, x2 represents the word embedding of the label, label represents positive and negative samples, the positive sample is 1, the negative sample is -1, and margin represents a hyperparameter.
[0095] The similarity-based image retrieval method provided by the embodiments of this application. First, retrieve information is obtained; the retrieve information includes: retrieve text and retrieve image; then, the retrieve information is input into the feature extraction network corresponding to the retrieve information in the target model to obtain the target feature vector of the retrieve information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieve text, and an image feature extraction network corresponding to the retrieve image; finally, multiple candidate feature vectors whose similarity to the target feature vector meets a preset similarity are selected from the vector library, and each image corresponding to each candidate feature vector among the multiple candidate feature vectors is used as the retrieve result image corresponding to the retrieve information; among them, the text feature extraction network includes: a text feature extraction model, and a feed-forward neural network; the feed-forward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a maximum pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image. In this way, the model is trained through the feature similarity between the image and the text, so that the model can perform image search by text and image search by image based on the feature similarity.
[0096] It should be noted that for the similarity-based image retrieval method provided by the embodiments of this application, the execution subject can be a similarity-based image retrieval device, or a control module in the similarity-based image retrieval device for executing the similarity-based image retrieval method. In the embodiments of this application, the similarity-based image retrieval method executed by the similarity-based image retrieval device is taken as an example to illustrate the similarity-based image retrieval device provided by the embodiments of this application.
[0097] It should be noted that in the embodiments of this application, the similarity-based image retrieval methods shown in the above respective method drawings are all exemplarily illustrated by taking one drawing in the embodiments of this application as an example. Specifically, when implemented, the similarity-based image retrieval methods shown in the above respective method drawings can also be implemented in combination with any other combinable drawings schemed in the above embodiments, which will not be elaborated here.
[0098] The image retrieval device based on similarity provided by the present application will be described below. The following description can be correspondingly referred to the image retrieval method based on similarity described above.
[0099] Figure 5 The structural schematic diagram of the image retrieval device based on similarity provided by an embodiment of the present application is shown in Figure 5 and specifically includes:
[0100] An acquisition module 501, configured to acquire retrieval information; the retrieval information includes: retrieval text and retrieval image; a feature extraction module 502, configured to input the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; a feature matching module 503, configured to screen out multiple candidate feature vectors from a vector library whose similarity to the target feature vector meets a preset similarity, and use the image corresponding to each candidate feature vector among the multiple candidate feature vectors as the retrieval result image corresponding to the retrieval information; wherein, the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is configured to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, a max pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image.
[0101] Optionally, the feature extraction module 502 is specifically configured to input the retrieval text into the text feature extraction model for feature extraction to obtain a first feature vector; the feature extraction module 502 is further specifically configured to input the first feature vector into the feedforward neural network for feature dimension adjustment to obtain a second feature vector, and determine the second feature vector as the target feature vector; wherein, the dimension of the second feature vector is the same as the dimension of the image feature vector output by the image feature extraction network; the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is configured to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network.
[0102] Optionally, the feature extraction module 502 is specifically configured to input the retrieved image into the multiple feature extraction sub-networks, perform feature extraction in sequence, and input the features output by each feature extraction sub-network in the multiple feature extraction sub-networks into the residual block for processing to obtain the image vector feature of the retrieved image; the feature extraction module 502 is further specifically configured to determine the image vector feature of the retrieved image as the target feature vector; wherein, the image feature extraction network includes: a plurality of feature extraction sub-networks connected in series, and a residual block; each feature extraction sub-network in the plurality of feature extraction sub-networks connected in series includes: a plurality of convolutional layers matching the number of color channels of the input image, and a max pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of feature extraction sub-networks connected in series are all input into the residual block for processing to obtain the image vector feature of the input image.
[0103] Optionally, the feature extraction module 502 is specifically configured to input the multiple third feature vectors output by the convolutional layer in the target feature extraction sub-network into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature vector and the average pooling feature vector corresponding to the multiple third feature vectors; the feature extraction module 502 is further specifically configured to input the max pooling feature vector and the average pooling feature vector into the activation layer and the dropout layer, and output the image feature vector corresponding to the target feature extraction sub-network; wherein, the target feature extraction sub-network is any one of the multiple target feature extraction sub-networks.
[0104] Optionally, the feature matching module 503 is specifically configured to calculate the similarity values between the target feature vector and each of the to-be-matched feature vectors in the vector library, and determine the to-be-matched feature vectors with similarity values greater than the preset similarity threshold as the multiple candidate feature vectors; wherein, the similarity values between the target feature vector and each of the to-be-matched feature vectors in the vector library are calculated based on cosine similarity.
[0105] Optionally, the device further includes: a sample construction module; the sample construction module is configured to construct a plurality of positive samples in the sample that include an image and the description text corresponding to the image, and randomly sample the images and description texts of the plurality of positive samples to obtain a plurality of negative samples in the sample that include an image and the description text not corresponding to the image.
[0106] Optionally, the device further includes: a model training module; the model training module is configured to input the images in the target samples into the image feature extraction network of the target model, and input the description texts in the target samples into the text feature extraction network of the target model to train the target model; the model training module is further configured to adjust the model parameters of the target model based on the loss value of the cosine loss function.
[0107] The image retrieval device based on similarity provided by the present application first obtains retrieval information; the retrieval information includes: retrieval text and retrieval image; then, the retrieval information is input into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; finally, a plurality of candidate feature vectors whose similarity to the target feature vector meets a preset similarity are selected from the vector library, and the image corresponding to each candidate feature vector among the plurality of candidate feature vectors is used as the retrieval result image corresponding to the retrieval information; wherein, the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is configured to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of cascaded feature extraction sub-networks, and a residual block; each of the plurality of cascaded feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, a max pooling layer and an average pooling layer distributed in parallel; the features output by each of the plurality of cascaded feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image. In this way, the model is trained based on the feature similarity between the image and the text, so that the model can perform image search by text and image search by image based on the feature similarity.
[0108] Figure 6 An entity structure diagram of an electronic device is exemplified, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 may call logical instructions in the memory 630 to execute an image retrieval method based on similarity. The method includes: obtaining retrieval information; the retrieval information includes: retrieval text and a retrieval image; inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and images; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; screening out a plurality of candidate feature vectors from a vector library whose similarity to the target feature vector meets a preset similarity, and using the image corresponding to each candidate feature vector among the plurality of candidate feature vectors as the retrieval result image corresponding to the retrieval information; wherein, the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a max pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image.
[0109] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0110] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the similarity-based image retrieval method provided by each of the above methods. The method includes: obtaining retrieval information; the retrieval information includes: retrieval text and retrieval image; inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; screening out a plurality of candidate feature vectors from a vector library whose similarity to the target feature vector meets a preset similarity, and using each image corresponding to each candidate feature vector among the plurality of candidate feature vectors as a retrieval result image corresponding to the retrieval information; wherein, the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each of the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a maximum pooling layer and an average pooling layer distributed in parallel; features output by each of the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain an image vector feature of the input image.
[0111] In another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it executes the similarity-based image retrieval method provided by each of the above methods. The method includes: obtaining retrieval information; the retrieval information includes: retrieval text and retrieval image; inputting the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; screening out a plurality of candidate feature vectors from the vector library whose similarity to the target feature vector meets a preset similarity, and using the image corresponding to each candidate feature vector among the plurality of candidate feature vectors as the retrieval result image corresponding to the retrieval information; wherein, the text feature extraction network includes: a text feature extraction model, and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network; the image feature extraction network includes: a plurality of serially connected feature extraction sub-networks, and a residual block; each of the plurality of serially connected feature extraction sub-networks includes: a plurality of convolutional layers matching the number of color channels of the input image, and a max pooling layer and an average pooling layer distributed in parallel; the features output by each of the plurality of serially connected feature extraction sub-networks are input into the residual block for processing to obtain the image vector feature of the input image.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image retrieval method based on similarity, characterized in that, Including: Obtaining retrieval information; The retrieval information includes: retrieval text and retrieval image; Inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; wherein, the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; Screening out multiple candidate feature vectors from a vector library whose similarity with the target feature vector meets a preset similarity, and using each image corresponding to each candidate feature vector among the multiple candidate feature vectors as a retrieval result image corresponding to the retrieval information.
2. The method for similarity-based image retrieval according to claim 1, wherein The step of inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information includes: Inputting the retrieval text into the text feature extraction model for feature extraction to obtain a first feature vector; Inputting the first feature vector into the feedforward neural network for adjusting the feature dimension to obtain a second feature vector, and determining the second feature vector as the target feature vector; Wherein, the dimension of the second feature vector is the same as the dimension of the image feature vector output by the image feature extraction network; the text feature extraction network includes: a text feature extraction model and a feedforward neural network; the feedforward neural network is used to adjust the dimension of the text feature vector output by the text feature extraction model to be the same as the dimension of the image feature vector output by the image feature extraction network.
3. The method for similarity-based image retrieval according to claim 1, wherein The step of inputting the retrieval information into a feature extraction network corresponding to the retrieval information in a target model to obtain a target feature vector of the retrieval information includes: Inputting the retrieval image into the multiple feature extraction sub-networks for sequential feature extraction, and respectively inputting the features output by each feature extraction sub-network among the multiple feature extraction sub-networks into the residual block for processing to obtain an image vector feature of the retrieval image; Determining the image vector feature of the retrieval image as the target feature vector; Wherein, the image feature extraction network includes: a series of multiple feature extraction sub-networks and a residual block; each feature extraction sub-network in the series of multiple feature extraction sub-networks includes: multiple convolutional layers matching the number of color channels of the input image, and a maximum pooling layer and an average pooling layer distributed in parallel; the features output by each feature extraction sub-network in the series of multiple feature extraction sub-networks are respectively input into the residual block for processing to obtain an image vector feature of the input image.
4. The method for similarity-based image retrieval according to claim 3, wherein The step of inputting the retrieval image into the multiple feature extraction sub-networks for sequential feature extraction includes: Respectively inputting multiple third feature vectors output by the convolutional layer in the target feature extraction sub-network into the maximum pooling layer and the average pooling layer to obtain a maximum pooling feature vector and an average pooling feature vector corresponding to the multiple third feature vectors; Input the maximum pooling feature vector and the average pooling feature vector into the activation layer and the dropout layer, and output the image feature vector corresponding to the target feature extraction sub-network; Among them, the target feature extraction sub-network is any one of the multiple target feature extraction sub-networks.
5. The method for similarity-based image retrieval according to claim 1, wherein The screening of multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets the preset similarity includes: Calculate the similarity values between the target feature vector and each to-be-matched feature vector in the vector library, and determine the to-be-matched feature vectors with similarity values greater than the preset similarity threshold as the multiple candidate feature vectors; Among them, the similarity values between the target feature vector and each to-be-matched feature vector in the vector library are calculated based on cosine similarity.
6. The method for similarity-based image retrieval according to claim 1, wherein Before obtaining the target feature vector of the retrieval information by inputting the retrieval information into the feature extraction network corresponding to the retrieval information in the target model, the method further includes: Construct multiple positive samples including an image and a description text corresponding to the image in the sample, and randomly sample the images and description texts of the multiple positive samples to obtain multiple negative samples including an image and a description text not corresponding to the image in the sample.
7. The similarity-based image retrieval method according to claim 6, wherein After constructing multiple positive samples including an image and a description text corresponding to the image, and randomly sampling the images and description texts of the multiple positive samples to obtain multiple negative samples including an image and a description text not corresponding to the image, the method further includes: Input the image in the target sample into the image feature extraction network of the target model, and input the description text in the target sample into the text feature extraction network of the target model to train the target model; Adjust the model parameters of the target model based on the loss value of the cosine loss function.
8. An image retrieval device based on similarity, characterized in that The device includes: An acquisition module, configured to acquire retrieval information; the retrieval information includes: retrieval text and retrieval image; A feature extraction module, configured to input the retrieval information into the feature extraction network corresponding to the retrieval information in the target model to obtain the target feature vector of the retrieval information; the target model is trained based on the feature similarity between text and image; the target model includes: a text feature extraction network corresponding to the retrieval text, and an image feature extraction network corresponding to the retrieval image; A feature matching module, configured to screen out multiple candidate feature vectors from the vector library whose similarity to the target feature vector meets the preset similarity, and use the image corresponding to each candidate feature vector in the multiple candidate feature vectors as the retrieval result image corresponding to the retrieval information.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program runs to execute the similarity-based image retrieval method according to any one of claims 1 to 7.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the similarity-based image retrieval method according to any one of claims 1 to 7 through the computer program.