Image-based search method and apparatus, device, and storage medium
By adopting a combination method of self-supervised learning and indexing algorithms in the image search technology, the problem of time-consuming and cost-consuming manual labeling in the prior art is solved, and more efficient image feature extraction and matching search is achieved.
Patent Information
- Application Number
- PCT/CN2024/137071
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
The existing picture search technology requires a large amount of labeled data for model training, and manual labeling consumes time and cost, reducing technical efficiency.
Using a technical architecture based on self-supervised learning, a pre-trained feature extractor is used to extract features of the user input images, and the feature vectors are processed through the indexing algorithm to achieve fast matching search, avoiding the steps of manual annotation.
It improves the technical efficiency of searching pictures with pictures, saves time and cost, and enhances the accuracy and generalization ability of feature extraction.
Smart Images

Figure CN2024137071_12062025_PF_FP_ABST
Abstract
Description
Image-based search method, device, equipment, and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311660389.8, filed on December 5, 2023, entitled “Image-based search method, device, equipment and storage medium,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of data processing, and in particular to an image-based search method, apparatus, device, and storage medium. Background Art
[0004] With the development of electronic information technology, more and more people are choosing to shop online. Shopping platforms can search for matching product information based on user-entered keywords and present this matching product information to users. However, searching by keyword can lead to insufficient search accuracy due to issues such as inaccurate keyword descriptions or semantic gaps. To improve search accuracy, image-based search can be used. That is, matching product information can be searched based on the image entered by the user. Image-based search is more intuitive and provides a better user experience than keyword search. Currently, image-based search mainly uses supervised learning to train a feature extraction model. The feature extraction model is used to extract image features for matching searches. However, training a feature extraction model using supervised learning requires a large amount of labeled data for model training. The labels of this data must be manually annotated, which is time-consuming and costly, reducing the technical efficiency of image-based search. Summary of the Invention
[0005] The embodiments of the present application provide an image-based search method, apparatus, device, and storage medium, which can improve the technical efficiency of image search.
[0006] In a first aspect, an embodiment of the present application provides an image-based search method, comprising: using a pre-trained feature extractor to perform feature extraction on an image input by a user to obtain a feature vector of an image to be searched, the feature extractor being based on a set of searched images as an unlabeled data set, and using self-supervised learning to obtain the image feature vector output by the feature extractor representing the underlying texture features and high-level semantic features of the image; using an indexing algorithm to process the feature vector of the image to be searched to obtain an index vector of the image to be searched; determining an image index vector that matches the index vector of the image to be searched based on the index vector of the image to be searched and a preset index vector library, the image index vector in the index vector library being obtained based on the indexing algorithm and the image feature vector extracted by the feature extractor from the set of searched images; and feeding back the searched image corresponding to the matching image index vector based on the matched image index vector.
[0007] In the second aspect, an embodiment of the present application provides an image-based search device, including: a feature extraction module, which is used to use a pre-trained feature extractor to extract features from an image input by a user to obtain a feature vector of an image to be searched, the feature extractor is based on a set of searched images as an unlabeled data set, and is obtained by self-supervised learning, and the image feature vector output by the feature extractor represents the underlying texture features and high-level semantic features of the image; an indexing module, which is used to use an indexing algorithm to process the feature vector of the image to be searched to obtain an index vector of the image to be searched; a search module, which is used to determine an image index vector that matches the index vector of the image to be searched based on the index vector of the image to be searched and a preset index vector library, the image index vectors in the index vector library are obtained based on the index algorithm and the image feature vectors extracted by the feature extractor from the set of searched images; a feedback module, which is used to feed back the searched image corresponding to the matching image index vector based on the matching image index vector.
[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the image-based search method of the first aspect is implemented.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the image-based search method of the first aspect is implemented.
[0010] The embodiments of the present application provide an image-based search method, apparatus, device, and storage medium. A pre-trained feature extractor can be used to extract features from an image input by a user to obtain a feature vector of the image to be searched. An index algorithm is used to obtain an image index vector of the image feature vector to be searched. A matching image index vector is searched in an index vector library including image index vectors of the image to be searched based on the image index vector to be searched. The searched image corresponding to the matching image index vector is a searched image in the searched image set that matches the image input by the user, thereby realizing image search. The feature extractor is obtained through self-supervised learning based on the searched image set. The searched image set serves as an unlabeled data set for self-supervised learning. There is no need for manual labeling of the searched images in the searched image set, which saves a lot of time and cost and improves the technical efficiency of image search. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] FIG1 is a flowchart of an image-based search method provided in one embodiment of the present application;
[0013] FIG2 is a schematic diagram of an example of a feature extractor structure provided in an embodiment of the present application;
[0014] FIG3 is a flowchart of an image-based search method provided in another embodiment of the present application;
[0015] FIG4 is a schematic diagram of an example of an index vector library provided in an embodiment of the present application;
[0016] FIG5 is a flowchart of an image-based search method provided by another embodiment of the present application;
[0017] FIG6 is a logic diagram of an example of a self-supervised learning process provided by an embodiment of the present application;
[0018] FIG7 is a logic diagram of an example of an image-based search process provided in an embodiment of the present application;
[0019] FIG8 is a schematic diagram of the structure of an image-based search device provided in one embodiment of the present application;
[0020] FIG9 is a schematic structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0021] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating examples of the present application. It should be noted that the acquisition, storage, use, processing, etc. of information and data in the embodiments of the present application are authorized by the user or relevant agencies and comply with the relevant provisions of national laws and regulations.
[0022] With the development of electronic information technology, more and more people are choosing to shop online. Shopping platforms can search for matching product information based on user-entered keywords and present this matching product information to users. However, searching by keyword can lead to insufficient search accuracy due to issues such as inaccurate keyword descriptions or semantic gaps. To improve search accuracy, image-based search can be used. That is, matching product information can be searched based on the image entered by the user. Image-based search is more intuitive and provides a better user experience than keyword search. Currently, image-based search mainly uses supervised learning to train a feature extraction model. The feature extraction model is used to extract image features for matching searches. However, training a feature extraction model using supervised learning requires a large amount of labeled data for model training. The labels of this data must be manually annotated, which is time-consuming and costly, reducing the technical efficiency of image-based search.
[0023] The present application provides an image-based search method, apparatus, device and storage medium, which can be applied to image search scenarios in online shopping malls or other scenarios that require image search. The present application proposes a technical architecture based on self-supervised learning to realize the function of image search, which can learn rich visual features from a large number of searched images without any annotation of the searched images, and obtain a feature extractor that can extract features from the image. The image feature vector of the image extracted by the feature extractor can reflect the underlying texture features and high-level semantic features of the image, and improve the accuracy of the image feature vector in representing the image. By indexing, the searched image that matches the image input by the user can be quickly found in a large number of searched images and fed back to the user, thereby realizing image search and improving the technical efficiency of image search.
[0024] The image-based search method, apparatus, device, and storage medium provided in this application are described below.
[0025] In a first aspect, the present application provides an image-based search method, which can be performed by an image-based search device, apparatus, etc., and is not limited herein. FIG1 is a flowchart of an image-based search method provided in one embodiment of the present application. As shown in FIG1 , the image-based search method may include steps S101 to S104.
[0026] In step S101, a pre-trained feature extractor is used to extract features from an image input by a user to obtain a feature vector of the image to be searched.
[0027] The image feature to be found is an image feature vector extracted by the feature extractor from the image input by the user. The feature extractor is obtained by self-supervised learning based on the searched image set which is an unlabeled data set. The searched image set includes the searched image. In an embodiment of the present application, it is expected to search for one or more searched images that match the image input by the user from the searched image set. The searched image set is an unlabeled data set for the feature extractor to perform self-supervised learning, that is, there is no need to label the searched images in the searched image set, and the searched image set is used as a sample for the feature extractor to perform self-supervised learning. The image feature vector output by the feature extractor includes the underlying texture features and high-level semantic features of the image, that is, the image feature vector output by the feature extractor can reflect the underlying texture features and high-level semantic features of the image, and combines the advantages of the underlying texture features and high-level semantic features. The content of the image feature vector output by the feature extractor is richer and more accurate.
[0028] The feature extractor may include a multi-layer structure, and the feature extractor may perform multi-layer processing on the input image to output an image feature vector. FIG2 is a schematic diagram of an example of a feature extractor structure provided in an embodiment of the present application. As shown in FIG2 , the feature extractor may include an input layer 21, a chunking layer 22, a downsampling layer 23, a global pooling layer 24, and a self-attention sub-model 25.
[0029] The input layer 21 may include a convolution layer and a maximum pooling layer. The input layer 21 is used to convert the input image into a feature map. For example, the input layer 21 may convert the RGB channel data of the input image into a feature map.
[0030] The block layer 22 includes a plurality of blocks 221, and the block layer in FIG2 includes four blocks 221. Each block 221 includes a residual block, and the number of residual blocks in different blocks 221 may be the same or different. For example, in the order from left to right in FIG2, the first block 221 may include three residual blocks, the second block 221 may include four residual blocks, the third block 221 may include six residual blocks, and the fourth block 221 may include three residual blocks. Different blocks may process feature maps in different dimensions and sizes, and the feature maps processed by the blocks may represent underlying texture features and high-level semantic features. For example, in the order from left to right in FIG2, the number of channels of the feature maps processed by the four blocks 221 gradually increases, and the width and height gradually decrease.
[0031] The downsampling layer 23 is used to increase the dimension and reduce the size of the feature map processed by the block 221, that is, to increase the number of channels and reduce the size of the feature map. The downsampling layer 23 may include a convolutional layer, a batch normalization layer, and a rectified linear unit (ReLU) activation layer. If the dimension of the feature map processed by the block 221 is high enough and the size is small enough, the downsampling layer 23 may not process the feature map processed by the block 221. The multiple feature maps output by the downsampling layer 23, which correspond one-to-one to the block 221, have the same dimension.
[0032] The global pooling layer 24 is used to perform a global pooling operation on the feature maps output by the downsampling layer 23. For example, the global pooling layer 24 can concatenate the feature maps a1, a2, a3, and a4 output by the downsampling layer 23 to obtain a concatenated feature map a5.
[0033] The self-attention sub-model 25 is used to adaptively adjust the weight distribution of the underlying texture features and high-level semantic features based on the spliced feature map, and output the image feature vector a7. The self-attention sub-model 25 can be implemented as a vision transformer (ViT), which is a deep learning model based on the Transformer architecture that can be used to process computer vision tasks. The self-attention sub-model 25 can adaptively adjust the weights of the underlying texture features and high-level semantic features in the output image feature vector a7 according to the image content and the requirements of the search task, better integrate the underlying texture features and high-level semantic features, so that the output image feature vector a7 can improve the accuracy and robustness of image retrieval, that is, image search.
[0034] For ease of understanding, the following takes the processing of an image of a specific size as an example for explanation. For example, the size of an image or feature map can be expressed as C×W×H, where C is channel data, which can also be regarded as dimension, W is width, and H is height; the size of the input image is 3×224×224, and the size of the feature map output after the input image passes through the input layer 21 is 64×56×56; the feature map of size 64×56×56 is processed by the first block 221 to obtain a feature map of size 256×56×56, the feature map of size 256×56×56 is processed by the second block 221 to obtain a feature map of size 512×28×28, and the feature map of size 512×28×28 is processed by the third block 221. After processing, a feature map of size 1024×14×14 is obtained, and the feature map of size 1024×14×14 is processed by the fourth block 221 to obtain a feature map of size 2048×28×28; the feature maps obtained by the first block 221, the second block 221, and the third block 221 are processed by the downsampling layer 23; the feature map output by the first block 221 can be processed by three layers of the downsampling layer 23, the first layer processing obtains a feature map of size 512×28×28, the feature map of size 512×28×28 is processed by the second layer to obtain a feature map of size 1024×14×14, and the feature map of size 102 The 4×14×14 feature map is processed by the third layer to obtain a feature map of size 2048×7×7; the feature map output by the second block 221 can be processed by the second layer of the downsampling layer 23, and the first layer is processed to obtain a feature map of size 1024×14×14, and the feature map of size 1024×14×14 is processed by the second layer to obtain a feature map of size 2048×7×7; the feature map output by the third block 221 can be processed by the first layer of the downsampling layer 23, and the first layer is processed to obtain a feature map of size 2048×7×7; the downsampling layer 23 may not process the feature map output by the fourth block 221; the downsampling layer The four feature maps a1, a2, a3, and a4, each with a dimension of 2048, output by 23 are processed by the global pooling layer 24 to form a spliced feature map a5. This spliced feature map a5 is then input into the self-attention sub-model 25 along with the category identifier a6. The self-attention sub-model 25 outputs an image feature vector a7, which has a dimension of 2048. It should be noted that the image feature vector a7 output by the self-attention sub-model 25 can represent both the underlying texture features and the high-level semantic features of the image, and the weighting of these features is learned by the feature extractor during self-supervised learning. If the image input to the feature extractor is a user-input image, the image feature vector output by the feature extractor is the image feature vector to be found.
[0035] In some examples, to facilitate the training process of self-supervised learning and enhance the feature extractor's ability to recognize different view transformations of the same image, the feature extractor may further include a projection layer. As shown in FIG2 , the projection layer 26 may be used to project the image feature vector a7 output by the self-attention sub-model 25 into a low-dimensional space, where the dimension of the low-dimensional space is lower than the dimension of the image feature vector. For example, the projection layer 26 may project the image feature vector a7 of dimension 2048 output by the self-attention sub-model 25 into a low-dimensional vector of dimension 128. The projection layer 26 is actually a multi-layer perceptron that amplifies the invariant features in the image by projecting the image representation vector into a low-dimensional space, thereby maximizing the feature extractor's ability to recognize different view transformations of the same image.
[0036] In step S102, the feature vector of the image to be searched is processed using an index algorithm to obtain an index vector of the image to be searched.
[0037] The image index vector to be searched is an index of the image feature vector to be searched. The indexing algorithm is not limited herein and may include, for example, a tree indexing algorithm, a hash indexing algorithm, a graph indexing algorithm, a quantized indexing algorithm, and the like. The dimension of the image index vector to be searched may be consistent with the dimension of the image feature vector to be searched.
[0038] In step S103, an image index vector that matches the image index vector to be searched is determined based on the image index vector to be searched and a preset index vector library.
[0039] The image index vector that matches the image to be searched is the index vector of the searched image that matches the user-input image. By matching the image to be searched index vector with the index vector library, the searched image that matches the user-input image in the searched image set can be determined. The image index vectors in the index vector library are derived based on an indexing algorithm and image feature vectors extracted from the searched image set by a feature extractor. The index vector library includes multiple image index vectors. Each searched image in the searched image set can be obtained through a feature extractor to obtain a corresponding image feature vector. The image feature vectors are processed using the indexing algorithm to obtain an image index vector for the image feature vector of the searched image.
[0040] The image index vector that matches the image index vector to be searched can be determined by the similarity between the image index vector to be searched and the image index vectors in the index vector library. In some examples, the similarity between the image index vector and each image index vector in the index vector library can be calculated, and the first N image index vectors in the index vector library ranked in descending order of similarity are determined as image indexes that match the image index vector to be searched, where N is a positive integer. Alternatively, based on the first N image index vectors in the index vector library ranked in descending order of similarity, the first N image index vectors can be used for reordering, such as calculating the average vector of the first N image index vectors and the image index vector to be searched, calculating the similarity between the average vector and each image index vector in the index vector library, and determining the first N image index vectors in the index vector library ranked in descending order of similarity as image indexes that match the image index vector to be searched. More times and more types of reordering processes can also be performed, which are not limited here. In the above embodiments, the similarity between vectors can be reflected by Euclidean distance or other parameters.
[0041] In step S104, based on the matched image index vector, the searched image corresponding to the matched image index vector is fed back.
[0042] The image index vectors are mapped one-to-one to the searched images, and matching image index vectors are obtained. The searched images corresponding to the matching image index vectors can then be fed back to the user. For example, in an online shopping mall, after a user enters an image, the image-based search method can be used to retrieve product images provided by the online mall that match the user's input image.
[0043] In an embodiment of the present application, a pre-trained feature extractor can be used to extract features from an image input by a user to obtain a feature vector of the image to be searched. An index algorithm can be used to obtain an image index vector of the image feature vector to be searched. A matching image index vector is searched in an index vector library including an image index vector of the searched image based on the image index vector to be searched. The searched image corresponding to the matching image index vector is the searched image in the searched image set that matches the image input by the user, thereby realizing image search. The feature extractor is obtained based on self-supervised learning of the searched image set. The searched image set serves as an unlabeled data set for self-supervised learning. There is no need to manually label the searched images in the searched image set. On the one hand, a lot of time and cost are saved and the technical efficiency of image search is improved. On the other hand, it can also avoid sensitive information issues that may be involved when labeling the searched images. The feature vector of the image to be searched output by the feature extractor can represent the underlying texture features and high-level semantic features of the image. The feature vector of the image to be searched that is a fusion of the underlying texture features and the high-level semantic features contains richer and more accurate content, and the generalization ability of the feature extractor is also better, which can ensure the processing performance of the feature extractor for various types of images and further improve the accuracy of image search.
[0044] In some embodiments, the above step S102 can be specifically refined as follows: using an indexing algorithm to establish an initial index vector for the image feature vector to be searched; performing dimensionality reduction processing on the initial index vector to obtain the image index vector to be searched. In some examples, the principal component analysis (PCA) algorithm can be used for dimensionality reduction processing. For example, the principal component analysis algorithm is used to reduce the dimension of the initial index vector with a dimension of 2048 to an image index vector to be searched with a dimension of 512. It should be noted that the method of establishing the image index vector in the index vector library is basically the same as the method of establishing the image index vector to be searched. The indexing algorithm can be used to establish an initial index vector for the image feature vector of the image to be searched, and then the initial index vector is subjected to dimensionality reduction processing to obtain the image index vector. For example, the image index vector to be searched and the image index vector can be obtained according to the following formulas (1) to (4): X index =HNSW(X f ) 2048 (1) x qindex =HNSW(x qf ) 2048 (2) X pca =PCA(X index ) 512 (3) x qpca=PCA(x qf ) 512 (4)
[0045] Among them, X index is the initial index vector corresponding to the image index vector of the searched image; x qindex is the initial index vector corresponding to the image feature vector to be found; HNSW stands for Hierarchical Navigable Small World (HNSW) algorithm, which is an indexing algorithm; X f is the image index vector of the image being searched; x qf is the feature vector of the image to be found; X pca is the image index vector of the image being searched; x qpca is the index vector of the image to be searched; PCA stands for principal component analysis, a dimensionality reduction algorithm; 2048 and 512 are both vector dimensions. Dimensionality reduction can further improve the accuracy of image search.
[0046] In some embodiments, a hierarchical indexing method may be used to establish an index vector library, that is, the index vector library may have a hierarchical structure. For example, a hierarchical navigation small-world algorithm may be used to establish the index vector library. The image feature vectors of the searched images in the searched image set may be preprocessed to divide the image feature vectors of the searched images into a number of clusters. The similarity within each cluster is calculated, and a graph structure is created based on the similarity within the clusters. Each node in the graph structure represents an image feature vector of the searched image. The graph structure is divided into hierarchies, and for each node in each layer, the node's neighbor nodes are determined based on a similarity parameter between nodes, such as the Euclidean distance, and the node is connected to its neighbor nodes, thereby forming an index vector library.
[0047] Correspondingly, matching the image index vector to be searched in the hierarchical index vector library requires a layer-by-layer search. FIG3 is a flowchart of an image-based search method provided by another embodiment of the present application. FIG3 differs from FIG1 in that step S103 in FIG1 can be specifically refined into steps S1031 and S1032 in FIG3.
[0048] In step S1031 , a search is performed from the top layer of the index vector library downwards layer by layer, and the image index vector with the highest similarity to the image index vector to be searched is determined in each layer and added to the search list.
[0049] Each layer in the index vector library includes multiple nodes, and the nodes are image index vectors of the image to be searched. The image index vector with the highest similarity to the image index vector to be searched can be determined at the top layer of the index vector library, and the image index vector with the highest similarity in the top layer can be added to the search list; the image index vector with the highest similarity to the image index vector to be searched can be determined at the second layer of the index vector library, and the image index vector with the highest similarity in the second layer can be added to the search list; and so on, until the image index vector with the highest similarity to the image index vector to be searched is determined at the last layer of the index vector library, and the image index vector with the highest similarity in the last layer can be added to the search list. If Euclidean distance is used to represent similarity, the image index vector with the smallest Euclidean distance to the image index vector to be searched can be determined as the image index vector with the highest similarity to the image index vector to be searched. For example, the Euclidean distance can be calculated according to the following formula (5):
[0050] Where ρ(A,B) is the Euclidean distance between vector A and vector B; A i is the i-th element in vector A; B i is the i-th element in vector B; m is the dimension of the vector.
[0051] For example, Figure 4 is a schematic diagram of an example of an index vector library provided in an embodiment of the present application. As shown in Figure 4, the index vector library includes a three-layer structure. Node b1 is the image index vector to be searched. At the top layer, node b2 with the smallest Euclidean distance to node b1 is obtained, and node b2 is added to the search list. At the second layer, node b3 with the smallest Euclidean distance to node b1 is obtained, and node b3 is added to the search list. At the third layer, node b4 with the smallest Euclidean distance to node b1 is obtained, and node b4 is added to the search list. The search list then includes node b2, node b3, and node b4.
[0052] In step S1032 , an image index vector that matches the image index vector to be searched is determined based on the first N image index vectors in the search list that are arranged in descending order of similarity.
[0053] N is a positive integer. The search list includes the image index vectors in each layer of the index vector library that have the highest similarity to the image index vector to be searched.
[0054] In some examples, the first N image index vectors in the search list that are ordered by similarity from high to low may be directly determined as image index vectors that match the image index vector to be found.
[0055] In other examples, the average value vector of the first N image index vectors in the search list, which are arranged in descending order according to similarity, and the image index vector to be searched can be calculated; matching is performed in the index vector library based on the average value vector, and the first N image index vectors, which are arranged in descending order according to similarity with the average value vector, are determined as image index vectors that match the image index vector to be searched. After obtaining the first N image index vectors in the search list, which are arranged in descending order according to similarity, the first N image index vectors can be used for a secondary query, with the average value vector as the input for the secondary query. The process of performing a secondary query using the average value vector can be shown in the following equations (6) to (8): K qe =query(Q qe ) (7) K qe ={K qe1 ,K qe2 ,…,K qeN} (8)
[0056] Among them, Q qe is the mean vector; K i is the i-th image index vector in the first N image index vectors in the search list arranged in descending order of similarity; Q is the image index vector to be searched; query represents the matching in the index vector library; K qe is the image index vector that matches the image index vector to be searched, which may include the image index vector K qe1 , K qe2 ,……,K qeN .
[0057] The average value vector is used to match the image index vector in the index vector library. The matching method can be the same as the matching method for the image index vector to be searched in the index vector library. Alternatively, the similarity between the average value vector and each image index vector in the index vector library can be directly calculated, and the first N image index vectors ranked in descending order of similarity to the average value vector are determined as the image index vectors that match the image index vector to be searched. In embodiments of the present application, further reordering can be performed based on the secondary query, such as by combining learning-based reordering methods, feedback-based reordering, etc., to improve the accuracy of image retrieval, i.e., image search.
[0058] In some embodiments, a feature extractor can be obtained by pre-training the image set to be searched as an unlabeled dataset for self-supervised learning, so that the feature extractor can be directly used in the subsequent image search process. Figure 5 is a flowchart of an image-based search method provided by another embodiment of the present application. The difference between Figure 5 and Figure 1 is that the image-based search method shown in Figure 5 can also include steps S105 to S107.
[0059] In step S105 , a residual network model trained using an open source labeled data set is obtained.
[0060] To facilitate feature extractor training, a residual network model can be trained using an open-source labeled dataset, i.e., pre-trained. The structure of the residual network model is consistent with that of the feature extractor in the above embodiment. In subsequent steps, this pre-trained residual network model is iteratively learned to update its model parameters, thereby obtaining a feature extractor.
[0061] In step S106 , positive sample pairs and negative samples of the searched image in the searched image set are obtained.
[0062] Positive sample pairs for a searched image may include the searched image itself and an image derived based on the searched image. Negative samples for a searched image may include images representing objects different from those represented by the searched image. Negative samples may come from within or outside the searched image set. A memory database may be provided to store negative samples.
[0063] In some examples, a positive sample pair includes a first positive sample and a second positive sample, and the negative sample is different from the positive sample in the positive sample pair. A searched image in the searched image set may be determined as the first positive sample; and data augmentation processing is performed on the searched image in the searched image set to obtain a second positive sample.
[0064] In step S107, the residual network model is iteratively learned according to the positive sample pairs, the residual network model, the negative samples and the preset momentum encoder until the residual network model meets the iteration end condition, and the residual network model that meets the iteration end condition is determined as the feature extractor.
[0065] A momentum encoder can be pre-set, and the momentum encoder is used to extract the feature vector of the negative sample. The residual network model is adjusted by the loss value of the feature vector obtained by processing the positive sample and the feature vector obtained by processing the negative sample, and multiple iterative learning is performed to obtain a feature extractor. Specifically, the positive sample is input into the residual network model to process the feature vector output by the residual network model; the negative sample is input into the momentum encoder for processing to obtain the feature vector output by the momentum encoder; the loss value is determined based on the feature vector output by the residual network model and the feature vector output by the momentum encoder. If the loss value does not meet the iteration end condition or the number of iterations does not meet the iteration end condition, the model parameters of the residual network model are adjusted, and iterative learning continues until the loss value meets the iteration end condition or the number of iterations meets the iteration end condition. The iteration end condition can be pre-set. For example, if the iteration end condition is related to the loss value, the iteration end condition includes that the loss value is less than or equal to the preset loss value. If the iteration end condition is related to the number of iterations, the iteration end condition includes that the number of iterations is equal to the preset number of iterations.
[0066] Figure 6 is a logical diagram of an example of a self-supervised learning process provided by an embodiment of the present application. As shown in Figure 6, a first positive sample and a second negative sample of the searched image can be obtained based on the searched image; the first positive sample is input into the residual network model to obtain feature vector 1; the second positive sample is input into the residual network model to obtain feature vector 2; a negative sample is obtained from the memory bank; the negative sample is input into the momentum encoder to obtain feature vector 3; a loss function is compared based on feature vector 1, feature vector 2 and feature vector 3 to obtain a loss value, and the loss value is gradient-backed; when the loss value does not meet the iteration end condition or the number of iterations does not meet the iteration end condition, the model parameters of the residual network model are adjusted, and the above process is repeated until the loss value meets the iteration end condition or the number of iterations meets the iteration end condition.
[0067] The results of image search using the feature extractor obtained by multiple iterations of the self-supervised learning method proposed in the above embodiment and some existing supervised learning models on the open source dataset Oxford5k and the open source dataset Paris6k are shown in Table 1 below:
[0068] Table 1
[0069] As can be seen from Table 1, the embodiment of the present application has a better effect in the mean average precision (mAP) of the two open source datasets than the supervised learning models such as NetVLAD, Faster R-CNN, SIAM-FV, Nonmetric, and DELF, and has a more obvious performance advantage.
[0070] The results of image search using the feature extractor obtained after multiple iterations of the self-supervised learning method proposed in the above embodiment and some existing self-supervised learning models are shown in Table 2 below on the open source dataset Oxford5k and the open source dataset Paris6k:
[0071] Table 2
[0072] As shown in Table 2, the solution in this application's embodiment achieves superior mean average precision (mAP) on the two open-source datasets compared to self-supervised learning models such as MoCo v2 and BYOL. In particular, on the Paris6k dataset, the solution in this application's embodiment leads the BYOL model by 5.2% in mAP, fully demonstrating the effectiveness of the image-based search method in this application's embodiment.
[0073] For ease of understanding, the following example illustrates the process of training a feature extractor and using the feature extractor to perform image search in an embodiment of the present application. FIG7 is a logic diagram of an example of an image-based search process provided in an embodiment of the present application. As shown in FIG7 , the image-based search process may include a pre-training portion, a self-supervised learning portion, and a search portion after obtaining the feature extractor.
[0074] In the pre-training phase, supervised learning of the residual network model is performed using a labeled dataset, which is an open-source dataset. A loss value is calculated using a multi-layer perceptron (MLP) and a loss function. Supervised learning is then iterated based on the loss value until the pre-training stopping condition is met. The model parameters of the residual network model are then copied to the encoder in the self-supervised learning phase. This means that the pre-trained residual network model serves as the encoder in the self-supervised learning phase.
[0075] In the self-supervised learning part, the unlabeled data set can be used to perform self-supervised learning on the encoder. The searched image set can be used as the unlabeled data set. Specifically, the searched images in the searched image set can be normalized first, and the size of each searched image can be adjusted to 3×224×224. The normalized searched image set can be recorded as X={x1,x2,…,x M}∈R M×C×W×H , where M is the number of searched images in the searched image set, C is the number of channels, which can also be regarded as the dimension, W is the width of the image, and H is the height of the image. The searched image set can be subjected to data enhancement processing to obtain the searched image set after data enhancement processing is x i The image obtained after data augmentation processing. The unlabeled data set may include the searched image set and the searched image set after data augmentation processing. The encoder, i.e., the pre-trained residual network model, is used to extract features from the searched images in the searched image set and the images in the searched image set after data augmentation processing to obtain a first positive sample set X including the first positive sample. f ={x f1 ,x f2 ,…,x fM} and a second positive sample set including the second positive sample x fi is x i The image feature vector obtained by encoder feature extraction, for The image feature vector obtained by encoder feature extraction. The negative sample image can be obtained from the memory bank, and the image feature vector of the negative sample image is obtained, that is, the negative sample The contrast loss function NCELoss can be used to calculate the loss value between positive samples (including the first positive sample and the second positive sample) and negative samples. The loss value can be obtained according to the following formula (9):
[0076] Where L is the loss value; x fk is the first positive sample; is the second positive sample; τ is the temperature hyperparameter; b is the size of the memory bank; is the i-th negative sample in the memory bank.
[0077] The loss value can be used to pass gradients back to update the encoder model parameters. After multiple rounds of iterative learning, the encoder that meets the iteration end conditions is used as the feature extractor.
[0078] In the search part, a feature extractor can be used to extract features from the normalized searched image in the searched image set to obtain the image feature vector of the searched image to form an image feature vector set. The image feature vector set can be expressed as X f =Encoder(X) GeM , where Encoder represents the feature extractor, GeM represents the global pooling layer in the feature extractor adopts the GeM pooling method, and X f ={f1,f2,…,f N}∈R N×D Represents a set of image feature vectors, where D can be 2048, indicating a dimension of 2048.
[0079] The image input by the user can be normalized first, and the size of the image input by the user can be adjusted to 3×224×224. The normalized image can be recorded as x q ∈R 1×C×W×H The definitions of C, W and H can be found above and will not be repeated here. The normalized image is input into the feature extractor to obtain the image feature vector x to be searched output by the feature extractor. qf =Encoder(x q ) GeM The definitions of Encoder and GeM can be found above and will not be repeated here. D can be 2048, indicating a dimension of 2048.
[0080] The hierarchical navigation small world algorithm can be used to establish an index vector for the image feature vector of the searched image and the feature vector of the image to be searched. The principal component analysis algorithm is used to reduce the dimension of the initial index vector to obtain the image index vector X pca and the image index vector x to be searched qpca , please refer to the above formulas (1) to (4) for details, which will not be repeated here. Here, the dimension of 2048 can be reduced to the dimension of 512.
[0081] The Euclidean distance can be used to represent the similarity between vectors. The index vector x of the image to be searched can be calculated qpca With X pca ={f pca1 ,f pca2 ,…,f pcaM}. Get the Euclidean distance of each image index vector in the image to be searched x. qpca The first N image index vectors with the closest Euclidean distance to the image index vector to be searched can be recorded as K = {K1, K2, ..., K k The calculation of Euclidean distance can be shown as follows (10):
[0082] Among them, j∈M, 512 is the dimension.
[0083] The first N image index vectors K={K1, K2, ..., K k} and the index vector x to be searched qpca The features are summed and averaged to obtain an average vector. The average vector is then used to perform a query using the Euclidean distance, which is a re-ranking process. This process can be seen in equations (6) to (8) above and will not be repeated here. The searched images corresponding to the first N image index vectors obtained by re-ranking are fed back to the user as search results.
[0084] A second aspect of the present application provides an image-based search device. FIG8 is a schematic diagram of the structure of an image-based search device provided in an embodiment of the present application. As shown in FIG8 , the image-based search device 300 may include a feature extraction module 301, an indexing module 302, a search module 303, and a feedback module 304.
[0085] The feature extraction module 301 may be used to perform feature extraction on an image input by a user using a pre-trained feature extractor to obtain a feature vector of the image to be searched.
[0086] The feature extractor is based on the search image set as an unlabeled dataset and is obtained through self-supervised learning. The image feature vector output by the feature extractor represents the low-level texture features and high-level semantic features of the image.
[0087] The indexing module 302 may be configured to process the feature vector of the image to be searched using an indexing algorithm to obtain an index vector of the image to be searched.
[0088] The search module 303 may be configured to determine an image index vector that matches the image index vector to be searched based on the image index vector to be searched and a preset index vector library.
[0089] The image index vectors in the index vector library are obtained based on the index algorithm and the image feature vectors extracted by the feature extractor from the searched image set.
[0090] The feedback module 304 may be configured to feed back the searched image corresponding to the matched image index vector based on the matched image index vector.
[0091] In an embodiment of the present application, a pre-trained feature extractor can be used to extract features from an image input by a user to obtain a feature vector of the image to be searched. An index algorithm can be used to obtain an image index vector of the image feature vector to be searched. A matching image index vector is searched in an index vector library including an image index vector of the searched image based on the image index vector to be searched. The searched image corresponding to the matching image index vector is the searched image in the searched image set that matches the image input by the user, thereby realizing image search. The feature extractor is obtained based on self-supervised learning of the searched image set. The searched image set serves as an unlabeled data set for self-supervised learning. There is no need to manually label the searched images in the searched image set. On the one hand, a lot of time and cost are saved and the technical efficiency of image search is improved. On the other hand, it can also avoid sensitive information issues that may be involved when labeling the searched images. The feature vector of the image to be searched output by the feature extractor can represent the underlying texture features and high-level semantic features of the image. The feature vector of the image to be searched that is a fusion of the underlying texture features and the high-level semantic features contains richer and more accurate content, and the generalization ability of the feature extractor is also better, which can ensure the processing performance of the feature extractor for various types of images and further improve the accuracy of image search.
[0092] In some embodiments, the indexing module 302 may be specifically configured to: establish an initial index vector for the image feature vector to be searched using an indexing algorithm; and perform dimensionality reduction processing on the initial index vector to obtain the image index vector to be searched.
[0093] In some embodiments, the index vector library is a hierarchical structure. The search module 303 can be specifically configured to: start from the top layer of the index vector library and search downward layer by layer, determine the image index vector with the highest similarity to the image index vector to be searched in each layer and add it to the search list; and determine the image index vector that matches the image index vector to be searched based on the first N image index vectors in the search list, which are sorted in descending order of similarity, where N is a positive integer.
[0094] In some examples, the search module 303 can be specifically used to: calculate the average value vector of the first N image index vectors in the search list arranged in descending order of similarity and the image index vector to be searched; match in the index vector library according to the average value vector, and determine the first N image index vectors arranged in descending order of similarity with the average value vector as the image index vector that matches the image index vector to be searched.
[0095] In some embodiments, the image-based search device 300 may further include a self-supervised learning module. The self-supervised learning module may be used to: obtain a residual network model trained using an open-source labeled dataset; obtain positive sample pairs and negative samples of the searched image in the searched image set; iteratively learn the residual network model based on the positive sample pairs, the residual network model, the negative samples, and a preset momentum encoder until the residual network model meets an iteration termination condition, and determine the residual network model that meets the iteration termination condition as the feature extractor.
[0096] In some examples, the positive sample pair includes a first positive sample and a second positive sample, and the negative sample is different from the positive sample in the positive sample pair. The self-supervised learning module can be specifically configured to: determine a searched image in the searched image set as the first positive sample; and perform data augmentation processing on the searched image in the searched image set to obtain the second positive sample.
[0097] In some examples, the self-supervised learning module can be specifically used to: process the positive samples input into the residual network model to obtain the feature vector output by the residual network model; input the negative samples into the momentum encoder for processing to obtain the feature vector output by the momentum encoder; determine the loss value based on the feature vector output by the residual network model and the feature vector output by the momentum encoder. If the loss value does not meet the iteration end condition or the number of iterations does not meet the iteration end condition, adjust the model parameters of the residual network model and continue iterative learning until the loss value meets the iteration end condition or the number of iterations meets the iteration end condition.
[0098] In some embodiments, the feature extractor may include an input layer, a block layer, a downsampling layer, a global pooling layer, and a self-attention sub-model. The input layer is used to convert the input image into a feature map; the block layer includes multiple blocks, each block includes a residual block, and different blocks process the feature map in different dimensions and sizes. The feature map after block processing represents the underlying texture features and high-level semantic features; the downsampling layer is used to increase the dimension and reduce the size of the feature map after block processing; the global pooling layer is used to perform a global pooling operation on the feature map output by the downsampling layer to obtain a spliced feature map; the self-attention sub-model is used to adaptively adjust the weight distribution of the underlying texture features and high-level semantic features based on the spliced feature map, and output an image feature vector.
[0099] In some examples, the feature extractor may further include a projection layer configured to project the image feature vector output by the self-attention sub-model into a lower-dimensional space having a lower dimension than the image feature vector.
[0100] A third aspect of the present application provides an electronic device. FIG9 is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. As shown in FIG9 , the electronic device 400 includes a memory 401 , a processor 402 , and a computer program stored in the memory 401 and executable on the processor 402 .
[0101] In some examples, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0102] The memory 401 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the image-based search method according to the embodiments of the present application.
[0103] The processor 402 reads the executable program code stored in the memory 401 to run a computer program corresponding to the executable program code, so as to implement the image-based search method in the above embodiment.
[0104] In some examples, the electronic device 400 may further include a communication interface 403 and a bus 404. As shown in FIG9 , the memory 401, the processor 402, and the communication interface 403 are connected via the bus 404 and communicate with each other.
[0105] The communication interface 403 is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. Input devices and / or output devices can also be connected through the communication interface 403.
[0106] The bus 404 includes hardware, software, or both that couples the components of the electronic device 400 to each other. By way of example, and not limitation, the bus 404 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Where appropriate, the bus 404 may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.
[0107] In a fourth aspect, the present application further provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the image-based search method in the above-mentioned embodiment can be implemented, and the same technical effect can be achieved. To avoid repetition, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which is not limited here.
[0108] An embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the image-based search method in the above embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0109] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiment. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps after understanding the spirit of this application. In addition, for the sake of brevity, a detailed description of known method technologies is omitted here.
[0110] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.
[0111] Those skilled in the art should understand that the above embodiments are illustrative rather than restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, the specification and the claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other devices or steps; the quantifier "one" does not exclude a plurality; the terms "first" and "second" are used to identify names rather than to indicate any specific order. Any figure marks in the claims should not be understood as limiting the scope of protection. The functions of multiple parts appearing in the claims can be implemented by a separate hardware or software module. The fact that certain technical features appear in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.
Claims
1. An image-based search method, comprising: Using a pre-trained feature extractor to extract features from an image input by a user, and obtaining a feature vector of the image to be searched, wherein the feature extractor is based on a set of images to be searched as an unlabeled data set and is obtained by self-supervised learning, and the image feature vector output by the feature extractor represents the underlying texture features and high-level semantic features of the image; Processing the feature vector of the image to be searched by using an index algorithm to obtain an index vector of the image to be searched; Determine an image index vector that matches the image index vector to be searched according to the image index vector to be searched and a preset index vector library, wherein the image index vectors in the index vector library are obtained based on the index algorithm and the image feature vectors extracted by the feature extractor from the searched image set; According to the matched image index vector, the searched image corresponding to the matched image index vector is fed back.
2. The method according to claim 1, wherein: The step of processing the feature vector of the image to be searched by using an index algorithm to obtain an index vector of the image to be searched includes: Using the index algorithm to establish an initial index vector for the image feature vector to be found; The initial index vector is subjected to dimensionality reduction processing to obtain the image index vector to be searched.
3. The method according to claim 1, wherein: The index vector library is a hierarchical structure. The step of determining an image index vector that matches the image index vector to be searched based on the image index vector to be searched and a preset index vector library includes: Starting from the top layer of the index vector library, searching downward layer by layer, determining in each layer the image index vector with the highest similarity to the image index vector to be found, and adding it to the search list; According to the first N image index vectors in the search list which are arranged in descending order according to similarity, an image index vector matching the image index vector to be found is determined, where N is a positive integer.
4. The method according to claim 3, wherein: The step of determining, based on the first N image index vectors in the search list that are arranged in descending order according to similarity, an image index vector that matches the image index vector to be searched comprises: Calculate the average vector of the first N image index vectors in the search list, which are arranged in descending order according to similarity, and the image index vector to be searched; Matching is performed in the index vector library according to the average value vector, and the first N image index vectors arranged in descending order of similarity with the average value vector are determined as image index vectors matching the image index vector to be searched.
5. The method according to claim 1, further comprising: Get the residual network model trained using an open source labeled data set; Obtaining positive sample pairs and negative samples of the searched images in the searched image set; According to the positive sample pair, the residual network model, the negative sample and the preset momentum encoder, the residual network model is iteratively learned until the residual network model meets the iteration end condition, and the residual network model that meets the iteration end condition is determined as the feature extractor.
6. The method according to claim 5, wherein: The positive sample pair includes a first positive sample and a second positive sample. The negative sample is different from the positive sample in the positive sample pair. Acquiring a positive sample of the searched image in the searched image set includes: Determine a searched image in the searched image set as a first positive sample; Data enhancement processing is performed on the searched image in the searched image set to obtain a second positive sample.
7. The method according to claim 5, wherein: The iterative learning of the residual network model according to the positive sample pair, the residual network model, the negative sample and the preset momentum encoder until the residual network model meets the iteration end condition includes: Processing the positive sample pairs into the residual network model to obtain a feature vector output by the residual network model; Inputting the negative sample into the momentum encoder for processing to obtain a feature vector output by the momentum encoder; The loss value is determined according to the feature vector output by the residual network model and the feature vector output by the momentum encoder. If the loss value does not meet the iteration end condition or the number of iterations does not meet the iteration end condition, the model parameters of the residual network model are adjusted, and iterative learning is continued until the loss value meets the iteration end condition or the number of iterations meets the iteration end condition.
8. The method according to any one of claims 1 to 7, wherein: The feature extractor includes an input layer, a chunking layer, a downsampling layer, a global pooling layer, and a self-attention sub-model; The input layer is used to convert the input image into a feature map; the block layer includes a plurality of blocks, each block includes a residual block, different blocks process the feature map in different dimensions and sizes, and the feature map after the block processing represents the underlying texture features and high-level semantic features; The downsampling layer is used to increase the dimension and reduce the size of the feature map after block processing; the global pooling layer is used to perform a global pooling operation on the feature map output by the downsampling layer to obtain a spliced feature map; the self-attention sub-model is used to adaptively adjust the weight distribution of the underlying texture features and the high-level semantic features based on the spliced feature map, and output an image feature vector.
9. The method according to claim 8, wherein: The feature extractor also includes a projection layer; the projection layer is used to project the image feature vector output by the self-attention sub-model into a low-dimensional space, and the dimension of the low-dimensional space is lower than the dimension of the image feature vector.
10. An image-based search device, comprising: A feature extraction module is used to extract features from an image input by a user using a pre-trained feature extractor to obtain a feature vector of the image to be searched, wherein the feature extractor is based on a set of images to be searched as an unlabeled data set and is obtained by self-supervised learning, and the image feature vector output by the feature extractor represents the underlying texture features and high-level semantic features of the image; An indexing module, used for processing the feature vector of the image to be searched by using an indexing algorithm to obtain an index vector of the image to be searched; A search module, configured to determine an image index vector matching the image index vector to be searched based on the image index vector to be searched and a preset index vector library, wherein the image index vectors in the index vector library are obtained based on the index algorithm and the image feature vectors extracted by the feature extractor from the searched image set; The feedback module is used to feed back the searched image corresponding to the matched image index vector according to the matched image index vector.
11. An electronic device, comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image-based search method according to any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the image-based search method as claimed in any one of claims 1 to 9.
Citation Information
Patent Citations
Searching method and system for image
CN112612913A
Image retrieval method, device and equipment and computer readable storage medium
CN114329006A
Image retrieval method and device, electronic equipment and computer readable storage medium
CN115344734A
Image searching method and device, computer readable storage medium and electronic equipment
CN115687670A
Image-based search method and device, equipment and storage medium
CN117743612A