An image retrieval method, apparatus, device and storage medium
By using the methods of global feature initial screening and local feature secondary screening in image retrieval, the problem of low image retrieval accuracy in the prior art is solved, and the search accuracy in background occlusion and similar background scenes is significantly improved.
Patent Information
- Application Number
- CN202110609456.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-01
AI Technical Summary
In the prior art, the image retrieval method based on global features and local features has low retrieval accuracy in background scenarios with severe background occlusion and highly similar background.
The method of two image search is adopted, firstly, the first image set is initially screened based on global features, and then the target local features are obtained by using the area detection method, and the second image search is performed to obtain the second image set.
Through the process of first global feature screening and then local feature secondary screening, the accuracy of image retrieval is effectively improved, especially in background occlusion and similar background scenes.
Smart Images

Figure CN113239225B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, especially to the field of intelligent decision-making technology, and particularly to an image retrieval method, device, equipment and storage medium. Background Art
[0002] Image retrieval technology based on deep learning can include global feature retrieval methods and local feature retrieval methods. Global feature retrieval methods mainly include netvlad, RMAC, GEM, etc. These methods retrieve images by extracting global features of images, resulting in low retrieval accuracy for background scenes with severely occluded backgrounds and highly similar backgrounds. Local feature retrieval methods mainly include delf, which retrieves images by extracting multiple local features of images. It focuses on local features and pays less attention to global features, and is prone to errors for images with locally identical regions. It can be seen that the image retrieval accuracy of these two image retrieval methods needs to be improved. Summary of the Invention
[0003] Embodiments of this application provide an image retrieval method, device, equipment and storage medium, which can improve image retrieval accuracy.
[0004] In a first aspect, embodiments of this application provide an image retrieval method, including:
[0005] Obtain a target image, where the target image does not include a target element;
[0006] Extract the global feature of the target image;
[0007] Use the global feature of the target image to perform a first image retrieval to obtain a first image set;
[0008] Adopt a region detection method to obtain the target local feature of the target image;
[0009] Perform a second image retrieval in the first image set according to the target local feature of the target image to obtain a second image set.
[0010] Optionally, the step of using the global feature of the target image to perform a first image retrieval to obtain a first image set includes:
[0011] Obtain the global features of each image in the image library;
[0012] Calculate the similarity between the target image and each image in the image library according to the global feature of the target image and the global features of each image in the image library;
[0013] Determine at least one image from the image library whose similarity to the target image is greater than or equal to a first preset value according to the calculated similarity;
[0014] Construct a first image set including the at least one image.
[0015] Optionally, the obtaining the target local feature of the target image by using a region detection method includes:
[0016] Obtain the feature map corresponding to the target image;
[0017] Perform region detection on the target image by using a region detection method to obtain the target region of the target image;
[0018] Determine the local feature corresponding to the target region in the feature map;
[0019] Determine the local feature corresponding to the target region in the feature map as the target local feature of the target image.
[0020] Optionally, the performing a second image retrieval in the first image set according to the target local feature of the target image to obtain a second image set includes:
[0021] Obtain the target local features of each image in the first image set;
[0022] Calculate the similarity between the target image and each image in the first image set according to the target local feature of the target image and the target local features of each image in the first image set;
[0023] Determine at least one image from the first image set whose similarity to the target image is greater than or equal to a second preset value according to the calculated similarity;
[0024] Construct a second image set including the at least one image.
[0025] Optionally, the performing a second image retrieval in the first image set according to the target local feature of the target image to obtain a second image set includes:
[0026] Encode and represent the target local feature of the target image by using a preset codebook to obtain the code of the target local feature of the target image;
[0027] Perform a second image retrieval in the first image set by using the code of the target local feature of the target image to obtain a second image set.
[0028] Optionally, the encoding of the target local feature of the target image is used to perform a second image retrieval in the first image set to obtain a second image set, including:
[0029] Obtain the encoding of the target local feature of each image in the first image set;
[0030] According to the encoding of the target local feature of the target image and the encoding of the target local feature of each image in the first image set, determine the type number or quantity of the same clustering center features between the target image and each image in the first image set;
[0031] According to the type number or quantity of the same clustering center features between the target image and each image in the first image set, screen out the second image set from the first image set.
[0032] Optionally, the encoding of the target local feature of the target image is used to perform a second image retrieval in the first image set to obtain a second image set, including:
[0033] Obtain the encoding of the target local feature of each image in the first image set;
[0034] According to the encoding of the target local feature of the target image and the encoding of the target local feature of each image in the first image set, determine the encoding of the target local features with the same position and the same clustering center features between the target image and each image in the first image set;
[0035] Calculate the similarity between the encoding of the target local features with the same position and the same clustering center features between the target image and each image in the first image set;
[0036] According to the calculated similarity, screen out the second image set from the first image set.
[0037] In a second aspect, an embodiment of the present application provides an image retrieval device, including:
[0038] An acquisition module, configured to acquire a target image, where the target image does not include a target element;
[0039] An extraction module, configured to extract the global feature of the target image;
[0040] A retrieval module, configured to perform a first image retrieval by using the global feature of the target image to obtain a first image set;
[0041] The acquisition module is further configured to acquire the target local feature of the target image by using a region detection method;
[0042] The retrieval module is further configured to perform a second image retrieval in the first image set according to the target local feature of the target image, so as to obtain a second image set.
[0043] In a third aspect, an embodiment of the present application provides an image retrieval device, including a processor and a memory, where the processor and the memory are connected to each other. The memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method described in the first aspect.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method described in the first aspect.
[0045] In summary, the image retrieval device can obtain a target image and extract the global feature of the target image. The image retrieval device can use the global feature of the target image to perform a first image retrieval to obtain a first image set, and adopt a region detection method to obtain the target local feature of the target image. Therefore, a second image retrieval is performed in the first image set according to the target local feature of the target image to obtain a second image set. Compared with the prior art of performing image retrieval solely using global features or solely using local features, the process of first performing a preliminary screening of images based on global features and then using target local features for secondary screening in this application can effectively improve the accuracy of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 is a schematic flowchart of an image retrieval method provided by an embodiment of the present application;
[0048] Figure 2 is a schematic flowchart of another image retrieval method provided by an embodiment of the present application;
[0049] Figure 3 is a schematic structural diagram of an image retrieval device provided by an embodiment of the present application;
[0050] Figure 4 is a schematic structural diagram of an image retrieval device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0052] Please refer to Figure 1 , which is a schematic flowchart of an image retrieval method provided by an embodiment of the present application. This image retrieval method can be applied to an image retrieval device, and the image retrieval device can be a server or an intelligent terminal, etc., which can implement the image retrieval function. Specifically, the method can include the following steps:
[0053] S101. Obtain a target image, where the target image does not include a target element.
[0054] Among them, the target element can be a human body image, an image of a certain animal, an image of a certain object, etc. In one embodiment, the target element can be an image that obscures the background image. The target image can be the background image.
[0055] In one embodiment, the image retrieval device can obtain the target image collected by the target device. Among them, the target device can be an electronic device with a camera or an electronic device with an image download function. The electronic device here can be an intelligent terminal such as a mobile phone, a laptop, or a desktop computer.
[0056] In one embodiment, the image retrieval device can obtain a first image, such as the first image collected by the target device. When the first image does not include the target element, the image retrieval device can determine the first image as the target image. When the first image includes the target element, the image retrieval device can perform matte processing on the target element included in the first image to obtain the first image with the matte removed as the target image. Performing matte processing on the target element can reduce the influence of the target element on the accuracy of image retrieval. Or, when the first image includes the target element, the image retrieval device can also add a mask to the target element included in the first image to obtain the first image with the mask added as the target image. Adding a mask to the target element can reduce the influence of the target element on the accuracy of image retrieval.
[0057] It should be noted that in addition to using matte processing or adding a mask to make an image not include the target element in the embodiments of the present application, other image processing methods can also be used to make an image not include the target element, which will not be listed one by one here.
[0058] In one embodiment, the image retrieval device may obtain a target video, obtain consecutive multiple frames of images of the target video, and determine a target image based on a second image among the multiple frames of images, where the second image is any one of the multiple frames of images. The image retrieval device may determine the second image as the target image when the second image does not include a target element. When the second image includes a target element, the image retrieval device may perform matte extraction processing on the target element included in the second image to obtain the matte-extracted second image as the target image. Here, the second image is an image among the multiple frames of images that includes the target element. Alternatively, when the second image does not include a target element, the image retrieval device may add a matte to the target element included in the second image to obtain the second image with the added matte as the target image. In one embodiment, the multiple frames of images may be consecutive multiple frames of images within a specified time range.
[0059] In one embodiment, the image retrieval device may read the target video from a database.
[0060] In one embodiment, the image retrieval device may obtain the target video collected by a target device.
[0061] In one embodiment, the target video may be a video of a target category. For example, when this application is used for financial fraud identification, the target category may be financial-related categories such as wealth management, investment, stocks, shopping, lending, etc. Here, in other application scenarios, the target category may also be other categories, which are not limited here.
[0062] In one embodiment, the target video may also be a video of a target user. For example, when this application is used for financial fraud identification, the target user may be a user with a financial fraud record or a user reported to have financial fraud suspicion; or, the target user may also be a user on the regulatory user list.
[0063] S102. Extract the global features of the target image.
[0064] In the embodiments of the present application, the image retrieval device may extract the global features of the target image through an image feature extraction model.
[0065] In one embodiment, the manner in which the image retrieval device extracts the global features of the target image may be as follows: The image retrieval device inputs the target image into a trained convolutional neural network, and the trained convolutional neural network performs feature extraction on the target image to obtain a feature map corresponding to the target image, and performs pooling processing and / or fully connected processing on the feature map corresponding to the target image to obtain the global features of the target image. Here, the convolutional neural network may be resnet, vgg, mobilenet, etc. In one embodiment, the image retrieval device may output the feature map corresponding to the target image through the last convolutional layer of the trained convolutional neural network.
[0066] For example, if the size of the convolution kernel of the last convolutional layer is 1*7*7*512, the image retrieval device may input the target image into the trained convolutional neural network. After being processed by the last convolutional layer of the trained convolutional neural network, a feature map corresponding to the target image is output, including 49 local features of 512 dimensions, and pooling processing or fully connected processing is performed on the 49 local features of 512 dimensions of the target image to obtain the global feature of the target image.
[0067] In one embodiment, the feature extraction method may include fully convolutional processing. Among them, fully convolutional processing is a process of processing through multiple convolutional layers and multiple pooling layers. Here, the connection method between the multiple pooling layers and the multiple convolutional layers may be: at least one convolutional layer is connected to one pooling layer, and at least one convolutional layer is connected after the pooling layer, and so on until the last convolutional layer is connected.
[0068] In one embodiment, the trained convolutional neural network can be obtained in the following manner: the image retrieval device acquires multiple sample images, and the sample images do not include the target element; the image retrieval device constructs a training set using the multiple sample images. The training set includes multiple sample image pairs, and the multiple sample image pairs include a first sample image pair and a second sample image pair. The first sample image pair includes a first sample image and a second sample image whose background is similar to that of the first sample image. The second sample image pair includes the first sample image and a third sample image whose background is not similar to that of the first sample image; the image retrieval device inputs the training set into the initial convolutional neural network to train the initial convolutional neural network to obtain the trained convolutional neural network. The sample image can be a sample image of the target scene, such as a sample image of the financial fraud scene. Among them, the sample image whose similarity to the first sample image is greater than or equal to the preset similarity can be the second sample image, and the image whose similarity to the first sample image is less than the preset similarity can be the third sample image.
[0069] In one embodiment, the process of the image retrieval device training an initial convolutional neural network model using a training set to obtain a trained convolutional neural network model is as follows: The initial convolutional neural network model is trained using the training set and the labels carried by each sample image pair in the training set indicating whether the sample images are similar to obtain a trained convolutional neural network model. In one embodiment, the image retrieval device can use the training set and the labels carried by each sample image pair in the training set indicating whether the sample images are similar, and combine with the method of metric learning to train the initial convolutional neural network model to obtain a trained convolutional neural network model. Metric learning can learn the similarity between two sample images. Through metric learning, the feature representation distances between positive examples (sample images with similar backgrounds) can be made closer, and the feature representation distances between negative examples (sample images with dissimilar backgrounds) can be made farther.
[0070] S103. Perform a first image retrieval using the global feature of the target image to obtain a first image set.
[0071] In the embodiments of the present application, the image retrieval device can use the global feature of the target image and adopt a fast nearest neighbor retrieval method to retrieve a first image set from the image library. By adopting this process, a small number of globally similar images with a high similarity ranking can be obtained. Among them, the fast nearest neighbor retrieval method can be hash retrieval, binary tree retrieval, etc. The image library can include multiple images that do not include the target element. Or, the image library can include multiple background images of the target scene, such as background images of financial fraud scenes. In one embodiment, the training set can cover the sample images of each financial fraud scene in multiple financial fraud scenes, and the image library can include background images of at least one financial fraud scene in multiple financial scenes.
[0072] In one embodiment, the way for the image retrieval device to perform a first image retrieval using the global feature of the target image to obtain a first image set can be: The image retrieval device obtains the global features of each image in the image library, and calculates the similarity between the target image and each image in the image library according to the global feature of the target image and the global features of each image in the image library; The image retrieval device determines at least one image in the image library whose similarity with the target image is greater than or equal to a first preset value according to the calculated similarity, and constructs a first image set including at least one image. By adopting the above process, a small number of images similar to the background to be retrieved can be found from the image library.
[0073] In one embodiment, when the image retrieval device calculates the similarity between the target image and each image in the image library according to the global features of the target image and the global features of each image in the image library, the image retrieval device may calculate the similarity between the global features of the target image and the global features of each image in the image library, and determine the similarity between the target image and each image in the image library according to the global features of the target image and the global features of each image in the image library.
[0074] In one embodiment, if each image in the image library already has global features (the global features can be stored in the image library, or in a database, or in the image retrieval device), the global features of each image can be directly obtained. If each image does not have global features yet, the global features of each image can be extracted by referring to the method of extracting the global features of the target image.
[0075] S104. Obtain the target local features of the target image by using a region detection method.
[0076] In the embodiment of the present application, the image retrieval device can obtain the feature map corresponding to the target image, and use a region detection algorithm, such as the MSER (Maximally Stable Extremal Regions) algorithm, to perform region detection on the target image to obtain the target region. Then, the image detection device can determine the local features corresponding to the target region in the feature map, and determine the local features corresponding to the target region in the feature map as the target local features of the target image. By using the region detection algorithm, regions without distinguishability in the background can be removed, such as a large amount of blank regions, such as a white wall.
[0077] In one embodiment, after obtaining the target region, the image retrieval device can obtain the position information of the target region in the target image, and obtain the position information of the local features of the target image in the target image. According to the position information of the target region in the target image and the position information of the local features of the target image in the target image, the target local features corresponding to the target region in the feature map are determined as the target local features of the target image. This process completes the screening of the local features and retains the distinguishable local features. For example, 20 local features can be screened out from 49 local features by using this process.
[0078] In one embodiment, if each image in the image library already has target local features (the target local features can be stored in the image library, or in a database, or in the image retrieval device), the target local features of each image can be directly obtained. If each image does not have target local features yet, the target local features of each image can be obtained by referring to the method of obtaining the target local features of the target image.
[0079] S105. Perform a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set.
[0080] In the embodiments of the present application, the manner in which the image retrieval device performs a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set may be as follows: The image retrieval device obtains the target local features of each image in the first image set, and calculates the similarity between the target image and each image in the first image set according to the target local features of the target image and the target local features of each image in the first image set. Then, the image retrieval device determines at least one image in the first image set whose similarity to the target image is greater than or equal to a second preset value according to the calculated similarity, and constructs a second image set including at least one image.
[0081] In one embodiment, in the process of calculating the similarity between the target image and each image in the first image set according to the target local features of the target image and the target local features of each image in the first image set, the image retrieval device may calculate the similarity between the target local features of the target image and the target local features of each image in the first image set, and determine the similarity between the target image and each image in the first image set according to the similarity between the target local features of the target image and the target local features of each image in the first image set.
[0082] It can be seen that Figure 1 In the shown embodiments, the image retrieval device may obtain a target image and extract the global features of the target image; the image retrieval device may perform a first image retrieval using the global features of the target image to obtain a first image set, and adopt a region detection method to obtain the target local features of the target image, so as to perform a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set. By adopting this process, the image retrieval accuracy can be improved.
[0083] Please refer to Figure 2 , which is a schematic flowchart of another image retrieval method provided by the embodiments of the present application. This method can be applied to the aforementioned image retrieval device. This method may include the following steps:
[0084] S201. Obtain a target image, where the target image does not include a target element.
[0085] S202. Extract the global features of the target image.
[0086] S203. Perform a first image retrieval using the global features of the target image to obtain a first image set.
[0087] S204. Use the region detection method to obtain the target local features of the target image.
[0088] Among them, steps S201 - S204 can refer to Figure 1 Steps S101 - S104 in the embodiment, which will not be elaborated here.
[0089] S205. Use the preset codebook to encode and represent the target local features of the target image, and obtain the encoding of the target local features of the target image.
[0090] S206. Use the encoding of the target local features of the target image to perform a second image retrieval in the first image set, and obtain a second image set.
[0091] In steps S205 - S206, the image retrieval device can use the preset codebook to encode and represent the target local features of the target image, obtain the encoding of the target local features of the target image, and use the encoding of the target local features of the target image to perform a second image retrieval in the first image set, and obtain a second image set.
[0092] Among them, the codebook can include multiple codewords, each codeword corresponds to a clustering center feature, and the clustering center features corresponding to each codeword are different. In one embodiment, the codebook can be obtained in the following way: The image retrieval device uses a clustering algorithm, such as the K - means algorithm, to perform clustering processing on a large number of images that do not include the target element, obtains multiple clustering center features, and constructs a codebook based on the multiple clustering center features. The large number of images that do not include the target element can be the background images of a large number of target scenarios, such as the background images of financial fraud scenarios. Among them, the clustering center feature is the feature corresponding to the clustering center. Or, the image retrieval device can also use a clustering algorithm to perform clustering processing on the aforementioned training set, obtain multiple clustering center features, and construct a codebook based on the multiple clustering center features.
[0093] Among them, the image retrieval device encodes and represents the target local features of the target image using a preset codebook. There are two ways to obtain the encoding of the target local features of the target image as follows: ① Determine the cluster center feature closest to the target local feature from the cluster center features corresponding to each codeword; use the codeword corresponding to the cluster center feature closest to the target local feature of the target image to encode and represent the target local feature of the target image, and obtain the encoding of the target local feature of the target image. ② Determine the cluster center feature closest to the target local feature of the target image from the cluster center features corresponding to each codeword; calculate the distance value between the target local feature of the target image and the cluster center feature closest to the target local feature of the target image; use the distance value between the target local feature of the target image and the cluster center feature closest to the target local feature of the target image to encode and represent the target local feature of the target image, and obtain the encoding of the target local feature of the target image.
[0094] In one embodiment, the way for the image retrieval device to perform a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain the second image set can be: The image retrieval device obtains the encodings of the target local features of each image in the first image set, and determines the type number or quantity of the same cluster center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encodings of the target local features of each image in the first image set, so as to screen out the second image set from the first image set according to the type number or quantity of the same cluster center features between the target image and each image in the first image set.
[0095] In one embodiment, in the process of the image retrieval device screening out the second image set from the first image set according to the type number of the same cluster center features between the target image and each image in the first image set, it can determine the similarity between the target image and each image in the first image set according to the type number of the same cluster center features between the target image and each image in the first image set, and screen out the second image set from the first image set according to the determined similarity. In one embodiment, the way for the image retrieval device to determine the similarity between the target image and each image in the first image set according to the type number of the same cluster center features between the target image and each image in the first image set can be: Determine the type number of the same cluster center features between the target image and each image in the first image set as the similarity between the target image and each image in the first image set.
[0096] In one embodiment, corresponding to the aforementioned first encoding acquisition method, the image retrieval device may determine the number of clustering center features corresponding to each codeword in the codebook for the target image and the number of clustering center features corresponding to each codeword in the codebook for each image in the first image set based on the encoding of the target local features of the target image and the encoding of the target local features of each image in the first image set, and construct the term frequency representation of the target image and the term frequency representation of each image in the first image set according to the determined numbers, so as to calculate the similarity between the target image and each image in the first image set based on the term frequency representation of the target image and the term frequency representation of each image in the first image set, and screen out the second image set from the first image set according to the calculated similarity. Among them, the term frequency representation of the target image may include the number of clustering center features corresponding to each codeword in the codebook for the target image, and the term frequency representation of each image in the first image set may include the number of clustering center features corresponding to each codeword in the codebook for that image. In one embodiment, the image retrieval device may count the number of types of the same clustering center features between the target image and each image in the first image set according to the term frequency representation of the target image and the term frequency representation of each image in the first image set, so as to calculate the similarity between the target image and each image in the first image set according to the number of types of the same clustering center features between the target image and each image in the first image set, and screen out the second image set from the first image set according to the calculated similarity. In one embodiment, the image retrieval device may directly calculate the similarity between the term frequency representation of the target image and the term frequency representation of each image in the first image set, determine the similarity between the target image and each image in the first image set according to the similarity between the term frequency representation of the target image and the term frequency representation of each image in the first image set, and screen out the second image set from the first image set according to the calculated similarity. Here, the image retrieval device may determine the similarity between the term frequency representation of the target image and the term frequency representation of each image in the first image set as the similarity between the target image and each image in the first image set. Here, the similarity between the term frequency representation of the target image and the term frequency representation of each image in the first image set may be calculated by calculating the Euclidean distance or cosine similarity, etc.
[0097] For example, assume that the word frequency representation of Image 1 is: [2, 3, 0, 0, 0], and the word frequency representation of Image 2 is: [6, 1, 6, 3, 0]. Among them, the 2 in the word frequency representation of Image 1 indicates that there are two local features in Image 1 with the same encoding, and the encoding is the codeword located in the first position in the codebook. Among them, the 3 in the word frequency representation of Image 1 indicates that there are three local features in Image 1 with the same encoding, and the encoding is the codeword located in the second position in the codebook. Among them, the 6 in the first position of the word frequency representation of Image 2 indicates that there are six local features in Image 2 with the same encoding, and the encoding is the codeword located in the first position in the codebook. Among them, the 1 in the word frequency representation of Image 2 indicates that there is one local feature in Image 2 whose encoding is the codeword located in the second position in the codebook. Among them, the 6 in the third position of the word frequency representation of Image 2 indicates that there are six local features in Image 2 with the same encoding, and the encoding is the codeword located in the third position in the codebook. Among them, the 3 in the word frequency representation of Image 2 indicates that there are three local features in Image 2 with the same encoding, and the encoding is the codeword located in the fourth position in the codebook. The image retrieval device can determine that the number of types of the same cluster center features between the two images is 2 according to the 2 in the first position and the 3 in the second position in the word frequency representation of Image 1, and the 6 in the first position and the 1 in the second position in the word frequency representation of Image 2. Therefore, it can be determined that the similarity between Image 1 and Image 2 is 2. Or, the image retrieval device can directly calculate the similarity between the word frequency representation of Image 1 and the word frequency representation of Image 2, and determine the similarity between the word frequency representation of Image 1 and the word frequency representation of Image 2 as the similarity between Image 1 and Image 2.
[0098] It should be noted that in the above process, a word frequency representation method similar to BOW is used to obtain the word frequency representation of an image for an image, and the local features of different images can be normalized into a feature representation of a unified dimension, and then the similarity is calculated. For example, in the above example, assume that Image 1 has 5 local features and Image 2 has 16 local features. The local features of Image 1 and the local features of Image 2 are both aggregated into the above word frequency representation of 5 dimensions (the same as the codebook size, and the codebook size refers to the number of cluster center features included in the codebook) using a word frequency representation method similar to BOW, and then the similarity is calculated.
[0099] In one embodiment, corresponding to the aforementioned second encoding acquisition method, the way for the image retrieval device to perform a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain the second image set may be as follows: The image retrieval device obtains the encodings of the target local features of each image in the first image set, and based on the encoding of the target local features of the target image and the encodings of the target local features of each image in the first image set, determines the encodings of the target local features that have the same position and the same clustering center features between the target image and each image in the first image set; The image retrieval device calculates the similarity between the encodings of the target local features that have the same position and the same clustering center features between the target image and each image in the first image set, and based on the calculated similarity, filters out the second image set from the first image set. Here, the similarity between the encodings of the target local features that have the same position and the same clustering center features between the target image and each image in the first image set can be calculated by calculating the Euclidean distance or cosine similarity, etc. In one embodiment, when there are multiple target local features that have the same position and the same clustering center features between the target image and each image in the first image set, the image retrieval device can calculate multiple similarities, obtain the cumulative value or average value of the multiple similarities, and then filter out the second image set from the first image set based on the cumulative value or average value of the multiple similarities. Or, the image retrieval device can also determine the score corresponding to the cumulative value or average value based on the cumulative value or average value of the multiple similarities, and filter out the second image set from the first image set based on the determined score.
[0100] For example, assume that the preset codebook is a codebook of size 65536, that is, a codebook corresponding to 65536 clustering center features. Assume that each codeword in the codebook corresponds to a clustering center feature dimension of 4. The clustering center features and the codewords corresponding to the clustering center features (the numbers before the colon are the codewords, and the numbers after the colon are the clustering center features) are as follows:
[0101] 1: 0.1, 0.1, 0.1, 0.1
[0102] 2: 0.2, 0.2, 0.2, 0.2 . .
[0103] 65536: 0.9, 0.9, 0.9, 0.9
[0104] Assume that the target image is image x, and the first image set includes image y. The target local features of image x and the codewords corresponding to the clustering center features closest to the distance target local features (the numbers before the colon are the codewords, and the numbers after the colon are the target local features) are as follows: 1: 0 0 0 0.1 3:0 0 0.1 0 5:0 0.1 0 0 7:0.1 0 0 0 7:0.1 0.1 0 0
[0110] The codewords corresponding to the target local features of image y and the cluster center features closest to the target local features (the numbers before the colon are the codewords, and the numbers after the colon are the target local features) are as follows: 1:0 0 0 0.2 3:0 0 0.2 0 4:0 0.2 0.2 0 4:0 0.2 0.2 0.2
[0115] The image retrieval device can determine that the target local features with the same position and the same cluster center features between image x and image y are the target local features corresponding to codeword 1 and the target local features corresponding to codeword 3. If the encoding of the target local feature corresponding to codeword 1 in image x (which is the distance value between the target local feature corresponding to codeword 1 in image x and the cluster center feature corresponding to codeword 1) and the encoding of the target local feature corresponding to codeword 3 in image x (which is the distance value between the target local feature corresponding to codeword 3 in image x and the cluster center feature corresponding to codeword 3) are r(x1) and r(x3) respectively, and the encoding of the target local feature corresponding to codeword 1 in image y (which is the distance value between the target local feature corresponding to codeword 1 in image y and the cluster center feature corresponding to codeword 1) and the encoding of the target local feature corresponding to codeword 3 in image y (which is the distance value between the target local feature corresponding to codeword 3 in image y and the cluster center feature corresponding to codeword 3) are r(y1) and r(y3) respectively, then the image retrieval device can calculate the similarities d(x1, y1) and d(x3, y3) between r(x1) and r(y1), and r(x3) and r(y3) respectively, and determine d(x1, y1) + d(x3, y3) as the similarity between image x and image y. Where:
[0116] r(x1) = [-0.1 -0.1 -0.1 0]
[0117] r(x3) = [-0.2 -0.2 -0.1 -0.2]
[0118] r(y1) = [-0.1 -0.1 -0.1 0.1]
[0119] r(y3) = [-0.2 -0.2 0 -0.2]
[0120] In one embodiment, the image retrieval device may also use a preset codebook to encode and represent the target local features of the target image, obtain the encoding of the target local features of the target image, and use the trained transformer network to extract target features based on the encoding of the target local features of the target image, and perform a second image retrieval in the first image set based on the target features. Among them, the process of performing retrieval in the first image set based on the target features may be: obtaining the target features of each image in the first image set, calculating the similarity between the extracted target features and each image in the first image set, and screening out a second image set from the first image set according to the calculated similarity.
[0121] It can be seen that Figure 2 In the embodiment shown, the image retrieval device may use a preset codebook to encode and represent the target local features of the target image, obtain the encoding of the target local features of the target image, and perform a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain a second image set, so as to realize a second image retrieval in the first image set according to the target local features of the target image and obtain a second image set. This process can improve the accuracy of image detection.
[0122] In an application scenario, since the methods and means of financial fraud such as online video lending vary, the image retrieval solution described in the embodiments of the present application can be used in scenarios where multiple people commit fraud at the same location and in the same scene, such as when a large number of business application locations are concentrated and the video backgrounds are the same during the application. The image retrieval solution described in the embodiments of the present application has the following advantages: when facing various harsh scenarios such as harsh lighting, motion blur, and packet loss in network transmission, which have high generalization requirements for the image retrieval process, the image retrieval solution provided in the embodiments of the present application also has high image retrieval accuracy for such harsh scenarios; when a large area of the human body is blocked, resulting in a relatively small proportion of useful background areas, the non-background areas can be automatically masked through the region detection algorithm to focus more on the background area, thereby realizing a more accurate image retrieval process; and, when facing an environment with many similar backgrounds, such as different in-vehicle environments, due to their high overall similarity, the embodiments of the present application can extract local background differences in the background for use in the image retrieval process; in addition, due to camera movement or people walking, there may be basically no overlapping areas in the originally similar background scenes, but the overall environment is similar. The embodiments of the present application can effectively focus on the global features of the image for use in the image retrieval process.
[0123] This application relates to blockchain technology. For example, the target image can be obtained from the blockchain, or the first image can be obtained from the blockchain, or the target video can also be obtained from the blockchain, and so on.
[0124] Please refer to Figure 3, which is a schematic structural diagram of an image retrieval device provided by an embodiment of the present application. This device can be applied to the aforementioned image retrieval device. Specifically, the device may include:
[0125] An acquisition module 301, configured to acquire a target image, where the target image does not include a target element.
[0126] An extraction module 302, configured to extract the global feature of the target image.
[0127] A retrieval module 303, configured to perform a first image retrieval by using the global feature of the target image to obtain a first image set.
[0128] The acquisition module 301 is further configured to acquire the target local feature of the target image by using a region detection method.
[0129] The retrieval module 303 is further configured to perform a second image retrieval in the first image set according to the target local feature of the target image to obtain a second image set.
[0130] In an alternative embodiment, the retrieval module 303 is specifically configured to:
[0131] Obtain the global features of each image in the image library;
[0132] Calculate the similarity between the target image and each image in the image library according to the global feature of the target image and the global features of each image in the image library;
[0133] Determine at least one image in the image library whose similarity with the target image is greater than or equal to a first preset value according to the calculated similarity;
[0134] Construct a first image set including the at least one image.
[0135] In an alternative embodiment, the acquisition module 301 is specifically configured to:
[0136] Obtain the feature map corresponding to the target image;
[0137] Perform region detection on the target image by using a region detection method to obtain the target region of the target image;
[0138] Determine the local feature corresponding to the target region in the feature map;
[0139] Determine the local feature corresponding to the target region in the feature map as the target local feature of the target image.
[0140] In an alternative embodiment, the retrieval module 303 is specifically configured to:
[0141] Obtain the target local features of each image in the first image set;
[0142] Calculate the similarity between the target image and each image in the first image set according to the target local features of the target image and the target local features of each image in the first image set;
[0143] Determine at least one image from the first image set whose similarity with the target image is greater than or equal to a second preset value according to the calculated similarity;
[0144] Construct a second image set including the at least one image.
[0145] In an alternative embodiment, the retrieval module 303 is specifically configured to:
[0146] Use a preset codebook to encode and represent the target local features of the target image to obtain the encoding of the target local features of the target image;
[0147] Perform a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain a second image set.
[0148] In an alternative embodiment, the retrieval module 303 is specifically configured to:
[0149] Obtain the encodings of the target local features of each image in the first image set;
[0150] Determine the type number or quantity of the same clustering center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encodings of the target local features of each image in the first image set;
[0151] Screen out a second image set from the first image set according to the type number or quantity of the same clustering center features between the target image and each image in the first image set.
[0152] In an alternative embodiment, the retrieval module 303 is specifically configured to:
[0153] Obtain the encodings of the target local features of each image in the first image set;
[0154] Determine the encodings of the target local features with the same position and the same clustering center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encodings of the target local features of each image in the first image set;
[0155] Calculate the similarity between the codes of the target local features with the same position and the same clustering center features between the target image and each image in the first image set;
[0156] According to the calculated similarity, screen out a second image set from the first image set.
[0157] It can be seen that Figure 3 In the illustrated embodiment, the image retrieval device can obtain a target image and extract the global features of the target image; the image retrieval device can use the global features of the target image to perform a first image retrieval to obtain a first image set, and adopt a region detection method to obtain the target local features of the target image, so as to perform a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set. By adopting this process, the image retrieval accuracy can be improved.
[0158] Please refer to Figure 4 , which is a schematic structural diagram of an image retrieval device provided by an embodiment of the present application. The image retrieval device described in this embodiment may include: one or more processors 1000 and a memory 2000. The processor 1000 and the memory 2000 can be connected through a bus.
[0159] The processor 1000 may be a central processing module (Central Processing Unit, CPU), and this processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), off-the-shelf programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0160] The memory 2000 may be a high-speed RAM memory or a non-volatile memory, such as a disk memory. Among them, the memory 2000 is used to store a computer program, and the computer program includes program instructions. The processor 1000 is configured to call the program instructions to execute the following steps:
[0161] Obtain a target image, and the target image does not include a target element;
[0162] Extract the global features of the target image;
[0163] Perform a first image retrieval using the global features of the target image to obtain a first image set;
[0164] Adopt a region detection method to obtain the target local features of the target image;
[0165] Perform a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set.
[0166] In one embodiment, when performing a first image retrieval using the global features of the target image to obtain a first image set, the processor 1000 is configured to call the program instructions and specifically execute the following steps:
[0167] Obtain the global features of each image in the image library;
[0168] Calculate the similarity between the target image and each image in the image library according to the global features of the target image and the global features of each image in the image library;
[0169] Determine at least one image in the image library whose similarity to the target image is greater than or equal to a first preset value according to the calculated similarity;
[0170] Construct a first image set including the at least one image.
[0171] In one embodiment, when adopting a region detection method to obtain the target local features of the target image, the processor 1000 is configured to call the program instructions and specifically execute the following steps:
[0172] Obtain the feature map corresponding to the target image;
[0173] Adopt a region detection method to perform region detection on the target image to obtain the target region of the target image;
[0174] Determine the local features corresponding to the target region in the feature map;
[0175] Determine the local features corresponding to the target region in the feature map as the target local features of the target image.
[0176] In one embodiment, when performing a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set, the processor 1000 is configured to call the program instructions and specifically execute the following steps:
[0177] Obtain the target local features of each image in the first image set;
[0178] Calculate the similarity between the target image and each image in the first image set according to the target local features of the target image and the target local features of each image in the first image set;
[0179] Determine at least one image from the first image set whose similarity to the target image is greater than or equal to a second preset value according to the calculated similarity;
[0180] Construct a second image set including the at least one image.
[0181] In one embodiment, when performing a second image retrieval in the first image set according to the target local features of the target image to obtain a second image set, the processor 1000 is configured to call the program instructions and specifically perform the following steps:
[0182] Encode and represent the target local features of the target image using a preset codebook to obtain the encoding of the target local features of the target image;
[0183] Perform a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain a second image set.
[0184] In one embodiment, when performing a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain a second image set, the processor 1000 is configured to call the program instructions and specifically perform the following steps:
[0185] Obtain the encoding of the target local features of each image in the first image set;
[0186] Determine the number or quantity of the types of the same clustering center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encoding of the target local features of each image in the first image set;
[0187] Screen out a second image set from the first image set according to the number or quantity of the types of the same clustering center features between the target image and each image in the first image set.
[0188] In one embodiment, when performing a second image retrieval in the first image set using the encoding of the target local features of the target image to obtain a second image set, the processor 1000 is configured to call the program instructions and specifically perform the following steps:
[0189] Obtain the encoding of the target local features of each image in the first image set;
[0190] Determine the encodings of the target local features that are in the same position and have the same clustering center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encodings of the target local features of each image in the first image set;
[0191] Calculate the similarity between the encodings of the target local features that are in the same position and have the same clustering center features between the target image and each image in the first image set;
[0192] Screen out a second image set from the first image set according to the calculated similarity.
[0193] In a specific implementation, the processor 1000 described in the embodiments of the present application may execute Figure 1 the embodiments, Figure 2 the implementation manners described in the embodiments, and may also execute the implementation manners described in the embodiments of the present application, which will not be elaborated herein.
[0194] In each embodiment of the present application, each functional module may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module.
[0195] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the computer-readable storage medium may be volatile or non-volatile. For example, the computer storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc. The computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application program required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0196] Among them, the blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0197] The above-disclosed is only a preferred embodiment of this application. Of course, the scope of rights of this application cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. An image retrieval method, characterized in that, it includes: Obtain a target image, where the target image does not include a target element; the target image is a background image, and the target element is an image content that obscures the background image; the way to obtain the target image is: obtain a first image; if the first image includes a target element, perform matte extraction on the target element included in the first image, and use the first image after matte extraction as the target image, or add a mask to the target element included in the first image, and use the first image with the mask added as the target image; Extract the global features of the target image; Use the global features of the target image to perform a first image retrieval to obtain a first image set; Adopt a region detection method to obtain the target local features of the target image; Use a preset codebook to encode and represent the target local features of the target image to obtain the encoding of the target local features of the target image; wherein, the codebook includes multiple codewords, each codeword corresponds to a clustering center feature, and the clustering center features corresponding to each codeword are different; the way to encode and represent the target local features includes: determine the clustering center feature closest to the target local features of the target image from the clustering center features corresponding to each codeword; calculate the distance value between the target local features of the target image and the clustering center feature closest to the target local features of the target image; use the distance value between the target local features of the target image and the clustering center feature closest to the target local features of the target image to encode and represent the target local features of the target image to obtain the encoding of the target local features of the target image; Use the encoding of the target local features of the target image to perform a second image retrieval in the first image set to obtain a second image set; the way of the second image retrieval includes: obtain the encoding of the target local features of each image in the first image set; according to the encoding of the target local features of the target image and the encoding of the target local features of each image in the first image set, determine the type number or quantity of the same clustering center features between the target image and each image in the first image set; according to the type number or quantity of the same clustering center features between the target image and each image in the first image set, screen out a second image set from the first image set.
2. The method according to claim 1, characterized in that, The step of using the global features of the target image to perform a first image retrieval to obtain a first image set includes: Obtain the global features of each image in the image library; According to the global features of the target image and the global features of each image in the image library, calculate the similarity between the target image and each image in the image library; According to the calculated similarity, determine at least one image in the image library whose similarity with the target image is greater than or equal to a first preset value; Construct a first image set including the at least one image.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining the target local feature of the target image by using the region detection method includes: Obtaining a feature map corresponding to the target image; Performing region detection on the target image by using the region detection method to obtain a target region of the target image; Determining a local feature corresponding to the target region in the feature map; Determining the local feature corresponding to the target region in the feature map as the target local feature of the target image.
4. The method according to claim 1, wherein, the step of performing a second image retrieval in the first image set by using the encoding of the target local feature of the target image to obtain a second image set includes: Obtaining the encodings of the target local features of each image in the first image set; Determining, according to the encoding of the target local feature of the target image and the encodings of the target local features of each image in the first image set, the encodings of the target local features with the same position and the same clustering center feature between the target image and each image in the first image set; Calculating the similarity between the encodings of the target local features with the same position and the same clustering center feature between the target image and each image in the first image set; Screening out a second image set from the first image set according to the calculated similarity.
5. An image retrieval device, wherein, it includes: An acquisition module, configured to acquire a target image, where the target image does not include a target element; the target image is a background image, and the target element is an image content that occludes the background image; the manner of acquiring the target image is: acquiring a first image; if the first image includes a target element, performing matte extraction on the target element included in the first image, and using the first image after matte extraction as the target image, or adding a mask to the target element included in the first image, and using the first image with the mask added as the target image; An extraction module, configured to extract the global feature of the target image; A retrieval module, configured to perform a first image retrieval by using the global feature of the target image to obtain a first image set; The acquisition module is further configured to obtain the target local feature of the target image by using a region detection method; The retrieval module is further configured to perform encoding representation on the target local feature of the target image by using a preset codebook to obtain an encoding of the target local feature of the target image; wherein, the codebook includes a plurality of codewords, each codeword corresponds to a clustering center feature, and the clustering center features corresponding to each codeword are different; the manner of performing encoding representation on the target local feature includes: determining, from the clustering center features corresponding to each codeword, the clustering center feature closest to the target local feature of the target image; calculating the distance value between the target local feature of the target image and the clustering center feature closest to the target local feature of the target image; using the distance value between the target local feature of the target image and the clustering center feature closest to the target local feature of the target image to perform encoding representation on the target local feature of the target image to obtain an encoding of the target local feature of the target image; The retrieval module is further configured to perform a second image retrieval in the first image set by using the encoding of the target local features of the target image, so as to obtain a second image set. The method for the second image retrieval includes: obtaining the encoding of the target local features of each image in the first image set; determining the type number or quantity of the same clustering center features between the target image and each image in the first image set according to the encoding of the target local features of the target image and the encoding of the target local features of each image in the first image set; and screening out the second image set from the first image set according to the type number or quantity of the same clustering center features between the target image and each image in the first image set.
6. An image retrieval device characterized in that it includes a processor and a memory, the processor and the memory are connected to each other. Wherein, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-4.
7. A computer-readable storage medium characterized in that the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-4.
Citation Information
Patent Citations
A remote sensing image retrieval method and system based on image segmentation and improved VLAD
CN109766467A
Image retrieval method and device and electronic equipment
CN110119460A