Image Retrieval Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

By updating the original image features so that they are in the same feature space as the target image features, and fusing the features to generate full image features, the problem of low image retrieval efficiency in the prior art is solved, and the saving of computing resources and the improvement of retrieval efficiency is achieved.

CN114329026BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111018523.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-01
Publication Date
2025-06-13
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

In the image retrieval, the prior art requires the feature extraction of new images and original images, resulting in waste of computing resources and reduced retrieval efficiency.

Method used

By acquiring the original image and adding the features of the target image, the original image features are updated, so that it is in the same feature space as the target image features, and the target image features are fused with the updated image features to generate a full number of image features for retrieval.

Benefits of technology

This method can quickly update image features, avoid re-extracting of original image features, save computing resources, and improve image retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329026B_ABST
    Figure CN114329026B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an image retrieval method, apparatus, electronic device, and computer-readable storage medium. After obtaining an original image, original image features corresponding to the original image, and at least one new target image, the embodiment of the present invention extracts features from the target image to obtain target image features corresponding to the target image. Then, the original image features are updated according to the target image features to obtain updated image features. Then, the target image features and the updated image features are fused to obtain all-image features for image retrieval. When an image retrieval request is received, features of the to-be-retrieved image carried in the image retrieval request are extracted, and based on the extracted to-be-retrieved image features and the all-image features, a target retrieved image is retrieved from the original image and the target image. This solution can improve the efficiency of image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to an image retrieval method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of Internet technologies, there is a huge amount of content on the Internet, such as a huge amount of images. Retrieving a required image from an image database requires comparison through image features. When there is an abnormality in image retrieval, it is often necessary to update the image data. When new images exist, existing image retrieval methods often need to uniformly extract features from the new images and the original images to generate a new image feature library, and then perform image retrieval based on the new image feature library.

[0003] In the process of researching and practicing the prior art, the inventors of the present invention found that uniformly extracting features from new images and original images means that the image features of the original images need to be extracted again, and the number of new images is often relatively small compared to the number of original images, resulting in a waste of computing resources and an increase in the extraction time of image features. Therefore, the efficiency of image retrieval is greatly reduced. Summary of the Invention

[0004] Embodiments of the present invention provide an image retrieval method, apparatus, electronic device, and computer-readable storage medium, which can greatly improve the efficiency of image retrieval.

[0005] An image retrieval method includes:

[0006] Obtaining an original image, the original image features corresponding to the original image, and at least one new target image, where the target image is an image newly added due to an error in retrieving the original image;

[0007] Extracting features from the target image to obtain target image features corresponding to the target image;

[0008] Updating the original image features according to the target image features to obtain updated image features, where the updated image features and the target image features are image features in the same feature space;

[0009] Fusing the target image features and the updated image features to obtain full-scale image features for image retrieval;

[0010] When receiving an image retrieval request, extracting features from the image to be retrieved carried in the image retrieval request, and retrieving a target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features.

[0011] Correspondingly, an embodiment of the present invention provides an image retrieval device, including:

[0012] An acquisition unit, configured to acquire an original image, the original image features corresponding to the original image, and at least one newly added target image, where the target image is an image newly added due to an incorrect retrieval of the original image;

[0013] An extraction unit, configured to extract features of the target image to obtain target image features corresponding to the target image;

[0014] An update unit, configured to update the original image features according to the target image features to obtain updated image features, where the updated image features and the target image features are image features in the same feature space;

[0015] A fusion unit, configured to fuse the target image features and the updated image features to obtain all-image features for image retrieval;

[0016] A retrieval unit, configured to, when receiving an image retrieval request, extract features of the to-be-retrieved image carried in the image retrieval request, and retrieve a target retrieval image from the original image and the target image according to the extracted to-be-retrieved image features and the all-image features.

[0017] Optionally, in some embodiments, the update unit may specifically be configured to obtain the original feature space information of the original image features and the target feature space information of the target image features, and compare the original feature space information and the target feature space information; based on the comparison result, use the feature mapping network of the trained image retrieval model to map the original image features to the feature space of the target image features to obtain updated image features.

[0018] Optionally, in some embodiments, the image retrieval device may further include a training unit, and the training unit may specifically be configured to obtain an image sample set, where the image sample set includes original image sample pairs and newly added target image sample pairs; extract image sample triples for training from the original image sample pairs and the target image sample pairs; train a first image retrieval network of a preset image retrieval model according to the original image sample pairs and the image sample triples to obtain a trained first image retrieval network; and jointly train the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples to obtain a trained image retrieval model.

[0019] Optionally, in some embodiments, the training unit may specifically be configured to obtain original image sample pairs, and collect images retrieved incorrectly by a second image retrieval network in the preset image retrieval model to obtain abnormal image sample pairs, where the second image retrieval network is trained from the original image sample pairs; obtain generalized images of the original image sample pairs to obtain generalized image sample pairs, and use the abnormal image sample pairs and the generalized image sample pairs as newly added target image sample pairs; fuse the target image sample pairs and the original image sample pairs to obtain the image sample set.

[0020] Optionally, in some embodiments, the training unit may specifically be configured to screen out target image samples from the original image sample pairs and the target image sample pairs; calculate the image distances between the target image samples and the remaining image samples in the image sample set respectively; based on the image distances, screen out a preset number of image negative samples corresponding to the target image samples in the image sample set, and add the image negative samples to the image sample pairs corresponding to the target image samples to obtain the image sample triples of the target image samples.

[0021] Optionally, in some embodiments, the training unit may specifically be configured to extract features of the original image samples in the original image sample pairs by using a first image retrieval network of the preset image retrieval model to obtain the original image sample features; screen out the target original image features corresponding to the image sample triples from the original image sample features; calculate the feature distances between the target original image features to obtain the original triple loss information corresponding to the image sample triples; converge the first image retrieval network based on the original triple loss information to obtain the trained first image retrieval network.

[0022] Optionally, in some embodiments, the training unit may specifically be configured to extract first image features of the image samples in the image sample set by using the trained first image retrieval network; extract features of the image samples in the image sample set by using the second image retrieval network, and map the extracted features by using a feature mapping network in the preset image retrieval model to obtain second image features; converge the trained first image retrieval network and the feature mapping network according to the first image features, the second image features, and the image sample triples to obtain the trained image retrieval model.

[0023] Optionally, in some embodiments, the training unit may specifically be configured to screen out the triples composed of the image samples in the target image sample pair from the image sample triples to obtain target image sample triples; screen out the triples composed of the image samples in the target image sample pair and the image samples in the original image sample pair from the image sample triples to obtain mixed image sample triples; determine the target loss information corresponding to the image sample set according to the target image sample triples, the mixed image sample triples, the first image feature, and the second image feature; and converge the trained first image retrieval network and the feature mapping network according to the target loss information to obtain a trained image retrieval model.

[0024] Optionally, in some embodiments, the training unit may specifically be configured to determine the triple loss information of the target image sample triples according to the first image feature to obtain target triple loss information; determine the triple loss information of the mixed image sample triples based on the second image feature to obtain mixed triple loss information; and fuse the target triple loss information and the mixed triple loss information to obtain the target loss information corresponding to the image sample set.

[0025] Optionally, in some embodiments, the training unit may specifically be configured to screen out any one image sample from the mixed image sample triples as the current image sample for calculating the triple loss information; calculate the triple loss information of the mixed image sample triples according to the current image sample and the second image feature to obtain initial mixed triple loss information; return to execute the step of screening out any one image sample from the mixed image sample triples as the current image sample for calculating the triple loss information until all the image samples in the mixed image sample triples are used as the current image sample to calculate the triple loss information, so as to obtain the initial mixed triple loss information corresponding to each image sample in the mixed image sample triples; and fuse the initial mixed triple loss information to obtain mixed triple loss information.

[0026] Optionally, in some embodiments, the retrieval unit may specifically be configured to calculate the feature distance between the image feature to be retrieved and each image feature in the full amount of image features to obtain a first feature distance; sort the first feature distances, and retrieve a target retrieval image from the original images and the target images according to the sorting result.

[0027] Optionally, in some embodiments, the retrieval unit may be specifically configured to update the cluster center corresponding to the original image according to the full-image features; based on the feature distance between the updated cluster center and the features of the image to be retrieved, filter out the target index from the index corresponding to the updated cluster center; and retrieve the target retrieved image from the original image and the target image according to the target index.

[0028] Optionally, in some embodiments, the retrieval unit may be specifically configured to filter out at least one candidate retrieved image corresponding to the target index from the original image and the target image; calculate the feature distance between the features of the image to be retrieved and the candidate retrieved image to obtain a second feature distance; and filter out the target retrieved image from the candidate retrieved images according to the second feature distance.

[0029] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the image retrieval method provided by the embodiment of the present invention.

[0030] In addition, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the image retrieval methods provided by the embodiment of the present invention.

[0031] After obtaining the original image, the original image features corresponding to the original image, and at least one new target image, the embodiment of the present invention extracts the target image features corresponding to the target image, then updates the original image features according to the target image features to obtain the updated image features, and then fuses the target image features and the updated image features to obtain the full-image features for image retrieval. When receiving an image retrieval request, the features of the image to be retrieved carried in the image retrieval request are extracted, and according to the extracted features of the image to be retrieved and the full-image features, the target retrieved image is retrieved from the original image and the target image; since this solution can migrate the original image features to the feature space of the target image features, the image features for image retrieval can be quickly updated without re-extracting the features of the original image, avoiding waste of computing resources, and also improving the extraction efficiency of the image features. Therefore, the efficiency of image retrieval can be improved. Description of the Drawings

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0033] Figure 1 It is a schematic diagram of the scenario of the image retrieval method provided by the embodiment of the present invention;

[0034] Figure 2 It is a schematic flowchart of the image retrieval method provided by the embodiment of the present invention;

[0035] Figure 3 It is a schematic diagram of the training process of the second image retrieval network provided by the embodiment of the present invention;

[0036] Figure 4 It is a network structure diagram of the joint training of the first image retrieval network and the feature mapping network provided by the embodiment of the present invention;

[0037] Figure 5 It is a schematic flowchart of the bucket retrieval provided by the embodiment of the present invention;

[0038] Figure 6 It is another schematic flowchart of the image retrieval method provided by the embodiment of the present invention;

[0039] Figure 7 It is a schematic structural diagram of the image retrieval device provided by the embodiment of the present invention;

[0040] Figure 8 It is another schematic structural diagram of the image retrieval device provided by the embodiment of the present invention;

[0041] Figure 9 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0043] The embodiments of the present invention provide an image retrieval method, device, and computer-readable storage medium. Among them, the image retrieval device can be integrated in an electronic device, and the electronic device can be a server or a terminal device, etc.

[0044] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart TV, a smart vehicle terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0045] For example, referring to Figure 1 , taking the case where the image retrieval device is integrated in an electronic device as an example, after the electronic device obtains the original image, the original image features corresponding to the original image, and at least one new target image, it extracts the features of the target image to obtain the target image features corresponding to the target image. Then, it updates the original image features according to the target image features to obtain the updated image features. Then, it fuses the target image features and the updated image features to obtain the full amount of image features for image retrieval. When receiving an image retrieval request, it extracts the features of the image to be retrieved carried in the image retrieval request, and retrieves the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full amount of image features, thereby improving the efficiency of image retrieval.

[0046] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0047] The image retrieval method of this embodiment can be executed by an electronic device, and the electronic device can be Figure 1 a server, or can be Figure 2 a terminal and other devices.

[0048] An image retrieval method includes:

[0049] Obtain an original image, the original image features corresponding to the original image, and at least one newly added target image. The target image is an image newly added due to an incorrect retrieval of the original image. Extract features from the target image to obtain the target image features corresponding to the target image. Update the original image features according to the target image features to obtain updated image features. The updated image features and the target image features are image features in the same feature space. Fuse the target image features and the updated image features to obtain the full-scale image features for image retrieval. When receiving an image retrieval request, extract features from the image to be retrieved carried in the image retrieval request, and retrieve the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features.

[0050] As Figure 2 shown, this image retrieval method can be executed by an electronic device, and the specific process is as follows:

[0051] 101. Obtain an original image, the original image features corresponding to the original image, and at least one newly added target image.

[0052] Among them, the target image is an image newly added due to an incorrect retrieval of the original image. The incorrect retrieval of the original image here can be understood as not retrieving the image that the user needs to retrieve in the original image or retrieving the wrong image. After this situation occurs, one or more target images need to be newly added to the image database composed of the original image. The role of newly adding the target image here is to make the image that the user needs to retrieve exist in the image database, thereby improving the accuracy of image retrieval.

[0053] Among them, there are various ways to obtain the original image, the original image features, and the target image. Specifically, it can be as follows:

[0054] For example, for the original image and the original image features, the original image and the original image features can be directly obtained. Or, the original image can also be obtained, and the original image retrieval model is used to extract features from the original image to obtain the original image features corresponding to the original image.

[0055] For a target image, the original image retrieval model can be used to extract features from the original image to obtain the original image features corresponding to the original image. The target image to be retrieved is input into the original image retrieval model, and the original image retrieval model retrieves candidate images from the original image according to the original image features. When the retrieved candidate image is abnormal or incorrect, the target image to be retrieved is used as the target image to be newly added. Alternatively, the image type of the original image can be obtained, and according to this image type, images different from this image type are screened out from the preset image database as the generalization images of the original image, and the generalization images are used as the target images. Alternatively, an image that needs to be newly added to the original image uploaded by the user through the terminal can also be received to obtain the target image.

[0056] 102. Extract features from the target image to obtain the target image features corresponding to the target image.

[0057] For example, the first image retrieval network of the trained image retrieval model can be used to extract features from the target image to obtain the target image features corresponding to the target image.

[0058] It should be noted that the first image retrieval network for extracting features from the target image is different from the image retrieval network for extracting the original image features of the original image. The original image features can be obtained by extracting features from the original image by the original image retrieval model in the image retrieval device or the second image retrieval network in the preset image retrieval model, or can also be obtained by extracting the original image by other devices or other equipment. The network structure of the first image retrieval network can be various. For example, it can be any network structure in the neural network (CNN). For example, it can be a residual network, a deep network or a neural network with other structures.

[0059] 103. Update the original image features according to the target image features to obtain the updated image features.

[0060] Among them, the updated image features and the target image features are image features in the same feature space. Therefore, this update can be understood as mapping the original image features to the feature space of the target image features, so that the target image features and the updated image features can perform feature retrieval with each other, and thus can form the image features in the image feature library of the image retrieval device.

[0061] Among them, there are various ways to update the original image features. Specifically, it can be as follows:

[0062] For example, the original feature space information of the original image features and the target feature space information of the target image features can be obtained, and the original feature space information and the target feature space information are compared. Based on the comparison result, the original image features are mapped to the feature space of the target image features by using the feature mapping network of the trained image retrieval model, and the updated image features are obtained.

[0063] Among them, based on the comparison result, there are various ways to map the original image features to the feature space of the target image features by using the feature mapping network. For example, based on the comparison result, the mapping parameters for mapping the original image features are determined, and based on the mapping parameters, the original image features are mapped to the feature space of the target image features by using the feature mapping network, so as to obtain the updated image features. Or, based on the comparison result, the mapping parameters corresponding to the comparison result are selected from the set of mapping parameters of the feature mapping network to obtain the target mapping parameters, and based on the target mapping parameters, the original image features are mapped to the feature space of the target image features by using the feature mapping network, so as to obtain the updated image features.

[0064] Among them, the trained image retrieval model can be set according to the requirements of actual applications. In addition, it should be noted that the trained image retrieval model can be set in advance by maintenance personnel or can be trained by the image retrieval device itself. That is, before the step of "based on the comparison result, the original image features are mapped to the feature space of the target image features by using the feature mapping network of the trained image retrieval model, and the updated image features are obtained", the image retrieval method may further include:

[0065] An image sample set is obtained. The image sample set includes original image sample pairs and newly added target image sample pairs. Image sample triples for training are extracted from the original image sample pairs and the target image sample pairs. The first image retrieval network of the preset image retrieval model is trained according to the original image sample pairs and the image sample triples to obtain the trained first image retrieval network. Based on the image sample set and the image sample triples, the trained first image retrieval network and the feature mapping network in the preset image retrieval model are jointly trained to obtain the trained image retrieval model.

[0066] Among them, there are various ways to obtain the image sample set. For example, original image samples can be obtained, and images retrieved incorrectly by the second image retrieval network in the preset image retrieval model are collected to obtain abnormal image sample pairs. The second image retrieval network is trained from the original image sample pairs, and the specific training process can be as Figure 3 shown. Generalized images of the original image sample pairs are obtained to obtain generalized image sample pairs. The abnormal image sample pairs and the generalized image sample pairs are used as the newly added target image sample pairs.

[0067] Among them, the original image sample pairs can be existing similar image pairs (i.e., image pairs marked whether two images are the same / similar), which become positive sample pairs. The original image sample pairs have been used to train the second image retrieval model. However, due to problems such as incomplete data coverage, the retrieval effect of some image styles is poor. The abnormal image sample pairs can be regarded as positive image sample pairs with the same / similar styles of retrieval error samples (badcases). All the collected abnormal image sample pairs can be considered as badcase samples under the second image retrieval network. The generalized image sample pairs can be samples from different domains from the original image sample pairs. For example, if there are more real-person scene samples in the original image sample pairs, the generalized domains can include positive sample pairs training sets such as game scenes and anime scenes.

[0068] After obtaining the original image samples, image sample triples for training can be extracted from the original image sample pairs and the target image sample pairs. The so-called triple can be an image sample group composed of an image sample, its corresponding positive sample, and negative sample. There are various ways to extract the prominent sample triples. For example, the target image samples can be selected from the original image sample pairs and the target image sample pairs, and the image distances between the target image samples and the remaining image samples in the image sample set are calculated respectively. Based on the image distances, a preset number of image negative samples corresponding to the target image samples are selected from the image sample set, and the image negative samples are added to the image sample pairs corresponding to the target image samples to obtain the image sample triples of the target image samples.

[0069] Among them, it can be found that the image sample triples are mainly triples mined from the similar samples of the image samples. For example, the image sample set can be divided into batches, with every bs image samples in a batch. In each batch (bs) of samples, any one image sample is selected as the target image sample x. The distance between it and x is calculated from the remaining (bs - 1) pairs of image samples (randomly selecting one image sample from each pair), sorted in ascending order of distance, and the TOP N samples are taken as the image negative samples, which are respectively combined with the image positive samples in x to form triples. It should be noted that in this solution, the learning is mainly for the same image samples. Therefore, for a certain image sample, all different image samples can be negative samples. The image samples sorted in ascending order above are sorted from similar to dissimilar relative to x. Since the difficult negative samples in the image sample triple learning have a better training effect on the preset image retrieval model, therefore, the most difficult N negative samples are selected as the image negative samples to form the image sample triples, and thus the image sample triples corresponding to each image sample in the original image sample pair and the target image sample pair can be obtained. Therefore, the composition of the image samples in the image sample triples can include three types: one is the original image sample pair + a certain image sample in other original image sample pairs, the second is the target image sample pair + a certain image sample in other target image sample pairs, and the third is the original image sample pair + a certain image sample in the target image sample pair / the target image sample pair + a certain image sample in the original image sample pair.

[0070] After extracting the image sample triples for training, the first retrieval network of the preset image retrieval model can be trained based on this triple to obtain the trained first image retrieval network. There are various training methods. For example, the first image retrieval network of the preset image retrieval model is used to extract features from the original image samples in the original image sample pair to obtain the original image sample features. The target original image features corresponding to the image sample triples are screened out from the original image sample features, the feature distances between the target original image features are calculated to obtain the original triple loss information corresponding to the image sample triples, and the first image retrieval network is converged based on the original triple loss information to obtain the trained first image retrieval network.

[0071] Among them, when training the first image retrieval network, the input is the full amount of original image samples. Therefore, the type of the image sample triple corresponding to the target original image features to be extracted can be a triple composed of an original image sample pair + a certain image sample in other original image sample pairs, which can also be called an original image sample triple. Therefore, there are various ways to screen out the target original image features corresponding to the original image sample triple in the original image features. For example, the image features corresponding to the target original sample in the original image sample triple can be screened out in the original image sample features to obtain the first original image features, the image features corresponding to the positive image sample of the target original image sample can be screened out in the original image sample features to obtain the second original image features, and the image features corresponding to the negative image sample of the target original image sample can be screened out in the original image features to obtain the third original image features. The first original image features, the second original image features, and the third original image features are used as the target original image features.

[0072] After screening out the target original image features, the feature distances between the target original image features can be calculated to obtain the original triple loss information corresponding to the image sample triple. There are various ways to calculate the original triple loss information. For example, the feature distance between the first original image feature and the second original image feature can be calculated to obtain the first feature distance, the feature distance between the first original image feature and the third original image feature can be calculated to obtain the second feature distance, and the distance difference between the first feature distance and the second feature distance can be calculated. The distance difference is fused with the preset margin distance to obtain the fused distance difference. The fused distance difference is compared with the preset distance threshold. When the fused distance difference exceeds the preset distance threshold, the fused distance difference is used as the original triple loss information. When the fused distance difference does not exceed the preset distance threshold, the preset distance threshold is used as the original triple loss information, which can be specifically shown in formula (1):

[0073] l tri =max(||x a -x p ||-‖x a -x n ‖+α,0) (1)

[0074] Among them, l tri is the original triple loss information, x a is the first original image feature, x p is the second original image feature, x n is the third original image feature, α is the preset margin distance, and the preset distance threshold is 0. When the image sample triple is an original image sample triple, α can be set to 0.8 or other values according to actual applications.

[0075] Among them, taking the first image retrieval network as a residual network (resnet101), taking this residual network as an example, the training process of the first image retrieval network can be as follows:

[0076] (1) Parameter initialization: Conv1 - Conv5 are initialized with the parameters of the second image retrieval network of the preset image retrieval model.

[0077] (2) Set learning parameters: Set the learning parameters as shown in Table 1 and Table 2, that is, all network parameters need to be learned.

[0078] (3) Learning rate: The basic feature and the image feature layer both adopt the learning rate of lr1 = 0.0005.

[0079] (4) Learning process: For the full amount of data, perform epoch (one round) iteration; process all samples once per round of iteration.

[0080] (5) The specific operations in each round of iteration are as follows: Divide every batch - size samples of the full amount of original image sample pairs into Nb batches. For each batch, perform the following training process:

[0081] A. Network forward: During training, the first image retrieval network performs forward calculation on an input original image sample to obtain a prediction result and outputs the original image feature;

[0082] B. Loss calculation: Find the original image sample triples in the batch samples and calculate the original triple loss information of the target original image features of the original image sample triples.

[0083] C. Update of the network parameters of the first image retrieval network: Adopt the SGD stochastic gradient descent algorithm or other gradient descent algorithms, perform gradient backward calculation on the original triple loss information to obtain the updated values of all network parameters, and update the first retrieval network.

[0084] D. Stopping condition (convergence condition): Record the average loss information of the first image retrieval network in each epoch. When the average loss information of the first image retrieval network has not decreased for 5 consecutive rounds (epoch), stop the training of the first image retrieval network.

[0085] Table 1 Structure table of ResNet - 101 feature module

[0086]

[0087] Table 2 Network structure of the first retrieval network

[0088] Network layer identifier Output size Network layer Pool 1x2048 Max pooling layer Embedding 1x128 Fully connected layer L2Norm 1x128 Normalization layer

[0089] After training the first image retrieval network, it is also possible to jointly train the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples. There are various ways of joint training. For example, the trained first image retrieval network can be used to extract features from the image samples in the image sample set to obtain the first image features, and the second image retrieval network can be used to extract features from the image samples in the image sample set to obtain the second image features. Based on the first image features, the second image features, and the image sample triples, the trained first image retrieval network and the feature mapping network are converged to obtain the trained image retrieval model.

[0090] Among them, there are various ways to converge the trained first image retrieval network and the feature mapping network based on the first image features, the second image features, and the image sample triples. For example, triples composed of the images in the target image sample pair can be selected from the image sample triples to obtain the target image sample triples, and triples composed of the images in the target image sample pair and the images in the original image sample pair can be selected from the image sample triples to obtain the mixed image sample triples. Based on the target image sample triples, the mixed image sample triples, the first image features, and the second image features, the target loss information corresponding to the image sample set is determined. Based on the target loss information, the trained first retrieval network and the feature mapping network are converged to obtain the trained image retrieval model.

[0091] Among them, the target image sample triples can be composed of three images in the target image sample pair, and the mixed image sample triples can be composed of the image pairs in the target image sample pair and the original image sample pair. For example, it can include the images in two target image sample pairs and the image in one original image sample pair, or it can also include the images in two original image sample pairs and the image in one target image sample pair. There are various ways to determine the target loss information corresponding to the image sample set based on the target image sample triples, the mixed image sample triples, the first image features, and the second image features. For example, based on the first image features, the triple loss information of the target image sample triples is determined to obtain the target triple loss information. Based on the second image features, the triple loss information of the mixed image sample triples is determined to obtain the mixed triple loss information. The target triple loss information and the mixed triple loss information are fused to obtain the target loss information corresponding to the image sample set.

[0092] Among them, the method of determining the target triple loss information according to the first image feature can refer to the calculation method of the original triple loss information, which will not be elaborated here. There are various ways to determine the mixed triple loss information according to the second image feature. For example, any one of the image samples in the mixed image sample triple can be selected as the current image sample for calculating the triple loss information. According to the current image sample and the second image feature, the triple loss information of the mixed image sample triple is calculated to obtain the initial mixed triple loss information. The step of selecting any one of the image samples in the mixed image sample triple as the current image sample for calculating the triple loss information is returned until all the image samples in the mixed image sample triple have been used as the current image sample to calculate the triple loss information, so as to obtain the initial mixed triple loss information corresponding to each image sample in the mixed image sample triple, and the initial mixed triple loss information is fused to obtain the mixed triple loss information.

[0093] Among them, the target image sample triple loss information is for the first image retrieval network branch, and the triple loss information of the target image sample triple (a-p-n) is calculated based on the first image feature. The mixed triple loss information is for the feature mapping network. For the mixed image sample triple (a-p-n), any one of them can be taken as the mapping feature (the second image feature mapped by the current image sample). For example, a can be taken as the mapping feature, or p can be taken as the mapping feature, or n can also be taken as the mapping feature. For example, for 10*bs triples, 30*bs mixed triplet losses can be obtained. The mixed triple loss information uses 0.6 as the margin (considering that the mixed features will actually cause a decrease in the similarity measurement effect. If the 0.8 margin is still used, it is easy to cause a large mixed triplet loss, making the proportion of the new feature triplet in the total loss smaller).

[0094] Among them, the method of calculating the triple loss information of the mixed image sample triple according to the current image sample and the second image feature to obtain the initial mixed triple loss information can refer to the calculation methods of the original triple loss information or the target triple loss information, which will not be elaborated one by one here.

[0095] Among them, there are various ways to fuse the target triple loss information and the mixed triple loss information to obtain the target loss information corresponding to the image sample set. For example, the weighted coefficients of the target triple loss information and the mixed triple loss information can be obtained, and according to the weighted coefficients, the target triple loss information and the mixed triple loss information are weighted respectively, and the weighted target triple loss information and the mixed triple loss information are fused to obtain the target loss information corresponding to the image sample set. Specifically, it can be shown in formula (2):

[0096] L = L new-triplet + 0.1 * L mix-triplet (2)

[0097] Among them, L is the target loss information, L new-triplet is the target triple loss information, and L mix-triplet is the mixed triple loss information. Considering that the update of the new feature (the first image feature) needs to give priority to ensuring the retrieval effect of the new feature, a small weight is used to weight the mapped feature. According to the principle of the above mixed triple loss information, the number of samples of the mixed triple loss information is three times that of the target triple loss information. The calculation formula of the target loss information can also be shown in formula (3):

[0098] L = L new-triplet + 0.1

[0099] * (L mix-triplet-a + L mix-triplet-p + L mix-triplet-n ) (3)

[0100] Among them, L is the target loss information, L new-triplet is the target triple loss information, L mix-triplet-a is the initial mixed triple loss information when the current image sample is a, and L mix-triplet-p is the initial mixed triple loss information when the current image sample you are p, and L mix-triplet-n is the initial mixed triple loss information when the current image sample is n.

[0101] When jointly training the first image retrieval network and the feature mapping network, the target triple loss information and the mixed triple loss information can be learned simultaneously. On the one hand, the old-new mapping of the features can be learned, and on the other hand, the new features can be fine-tuned so that they can support the retrieval of the mapped features.

[0102] Among them, the mapping structure of the feature mapping network is shown in Table 3, and it can be a FC_map mapping structure composed of one FC (fully connected) layer (of course, a multi-layer FC + relu activation module can be used as the mapping structure, which can be specifically determined according to data and model requirements). When the mapping structure is more complex, the inference of the mapping structure in applications is more time-consuming and resource-consuming. Therefore, the complexity of the mapping structure is also related to practical applications.

[0103] Table 3

[0104]

[0105]

[0106] The network structure for jointly training the first image retrieval network and the feature mapping network can be as Figure 4 shown. Taking the original image sample pair as the old image sample, the target image sample pair as the new image sample, the first image retrieval network as the new network, and the second image retrieval network as the old network as an example, the input sample passes through both the new and old networks simultaneously, generating new and old image features (embeddings). The old image features pass through the mapping module to generate mapped image features (the second image features). Using the target triplet loss information of the first image features and the mixed triplet loss information of the mapping structure and the new features as the supervision information, the feature mapping network and the new network (the first image retrieval network after training) are updated. The specific training process can be as follows:

[0107] 1) Parameter initialization: Use the learning result of the first new model (i.e., the first image retrieval network after training) as the initialization weight of the current training model.

[0108] 2) Set learning parameters: Set the learning parameters as in Table 1-2-3, that is, all network parameters need to be learned.

[0109] 3) Learning rate: The learning rate of the basic features and the embedding layer is both lr1 = 0.0005.

[0110] 4) Learning process: Perform epoch rounds of iteration on the full amount of data; process all samples once in each round of iteration;

[0111] 5) The specific operations in each round of iteration are as follows: Divide all samples (new and old image sample pairs) into batches of every batch-size samples, and divide them into Nb batches. For each batch, perform:

[0112] (1) Model forward: During training, the neural network performs forward calculation on an input image to obtain the prediction result and outputs the embedding result (the first image feature and the second image feature).

[0113] (2) Loss calculation: Calculate according to the above method for triplets in the batch samples. Depending on the composition of the triplet samples (the combination of badcase flags of the three images in the triplet), calculate the loss of a certain triplet, and sum up the loss information of all triplets to obtain the target loss information.

[0114] (3) Model parameter update: Use the SGD (Stochastic Gradient Descent) method (or other convergence algorithms) to perform backward gradient calculation on the target loss information in the previous step and obtain the updated values of all model parameters, and update the feature mapping network and the first image retrieval network after training.

[0115] 6) Stopping condition: Record the average loss information of the model in each epoch. When the average loss information of the model does not decrease for 5 consecutive rounds (epochs) or the preset number of rounds, stop the joint training of the first image retrieval network and the feature mapping network.

[0116] 104. Fuse the target image features and the updated image features to obtain the full-scale image features for image retrieval.

[0117] Among them, the full-scale image features can be all the image features for image retrieval.

[0118] Among them, there are various ways to fuse the target image features and the updated image features. Specifically, it can be as follows:

[0119] For example, based on the updated image features, update the preset image feature set to obtain the updated image feature set, and add the target image features to the updated image feature set to obtain the full-scale image features. Or, directly merge the target image features and the updated image features to obtain the full-scale image features.

[0120] 105. When receiving an image retrieval request, extract the features of the image to be retrieved carried in the image retrieval request, and retrieve the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features.

[0121] Among them, there are various ways to extract the features of the image to be retrieved carried in the image retrieval request. For example, the first image retrieval network in the trained image retrieval model can be used to extract the features of the image to be retrieved to obtain the features of the image to be retrieved of the image to be retrieved. Here, the first image retrieval network in the trained image retrieval model can be the first image retrieval network obtained after the joint training of the first image retrieval network after training.

[0122] After extracting the features of the image to be retrieved, the target retrieved image can be retrieved from the original image and the target image according to the features of the image to be retrieved and the features of all images. There are various retrieval methods. For example, the feature distance between the features of the image to be retrieved and each image feature in the features of all images can be calculated to obtain the first feature distance, the first feature distance can be sorted, and according to the sorting result, the target retrieved image can be retrieved from the original image and the target image. Or, the clustering center corresponding to the original image can be updated according to the features of all images, and based on the feature distance between the updated clustering center and the features of the image to be retrieved, the target index can be selected from the indexes corresponding to the updated distance center. According to the target index, the target retrieved image can be retrieved from the original image and the target image.

[0123] Among them, there are various ways to retrieve the target retrieved image from the original image and the target image according to the sorting result of the first feature distance. For example, the top K images in the sorting result can be selected from the original image and the target image as the target retrieved image, or the image with the smallest first feature distance can be selected from the original image and the target image as the target retrieved image. Or, the top K candidate images can be selected from the sorting result, the first feature distance corresponding to the candidate images can be compared with a preset distance threshold, and the images whose first feature distance exceeds the preset distance threshold can be selected from the candidate images to obtain at least one target image. The target images are classified, and according to the number and feature distance of each type of target image, at least one image is selected from the target images as the target image to be retrieved. For example, the image of the target type with the largest number is selected from the target images, and the top N images with the smallest feature distance are selected from the images of the target type as the target retrieved image.

[0124] Among them, according to the target index, at least one candidate retrieved image corresponding to the target index is selected from the original image and the target image, the feature distance between the features of the image to be retrieved and the features of the candidate retrieved image is calculated to obtain the second feature distance, and according to the second feature distance, the target retrieved image is selected from the candidate retrieved images.

[0125] Among them, the retrieval method using the index corresponding to the clustering center can be bucket retrieval based on clustering, such as Figure 5As shown, the original image features of the original image are extracted, the original image features are clustered to obtain at least one cluster center, the cluster center is used as the index for retrieval, and the association relationship between the index and the original image is established. At this time, when an image retrieval request is received, the image to be retrieved carried in the image retrieval request is extracted, the nearest index is found according to the features of the image to be retrieved, and the associated images of these indexes are obtained to get the candidate retrieval image retrieval recall. Calculate the feature distance between the image features of the retrieved image and the features of the image to be retrieved, sort them from small to large, and take the top K original images in the sorting as the retrieved images for output. When there is a new target image in the original image, the features of the target image are extracted, the original image features of the original image are updated, and the target image features and the updated image features are merged to obtain the full-scale image features. Then, the cluster center and the index corresponding to the cluster center are updated according to the full-scale features, so as to support optimized retrieval under partial updated features.

[0126] Optionally, after retrieving the target retrieval image corresponding to the image to be retrieved, the image to be retrieved can also be classified or recognized according to the category of the target retrieval image. Here, the classification or recognition mainly refers to the category-level recognition of the image to be retrieved, without considering the specific instance of the object, but only considering the category of the object (such as people, dogs, cats, birds, etc.) for recognition and giving the category to which the object belongs. There are various ways to classify or recognize. For example, when the number of target retrieval images is 1, the classification or recognition result of the image to be retrieved at this time can be the category of the target retrieval image. When the number of target retrieval images is multiple, the target retrieval images can be classified, and then, the number of images of each category of target retrieval images is obtained, and the category of the target image with the largest number of images is selected from the image categories as the classification or recognition result of the image to be retrieved. Or, when the number of target retrieval images is multiple, the target retrieval images are classified, and then, the total feature distance corresponding to each image category is calculated, and the image category with the smallest total feature distance is used as the classification or recognition result of the image to be retrieved.

[0127] As described above, after obtaining the original image, the original image features corresponding to the original image, and at least one new target image, the method extracts features from the target image to obtain the target image features corresponding to the target image. Then, the original image features are updated according to the target image features to obtain the updated image features. Then, the target image features and the updated image features are fused to obtain the full-scale image features for image retrieval. When an image retrieval request is received, features are extracted from the image to be retrieved carried in the image retrieval request, and the target retrieval image is retrieved from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features. Since this solution can transfer the original image features to the feature space of the target image features, the image features for image retrieval can be quickly updated without re-extracting the features of the original image, avoiding waste of computing resources and improving the extraction efficiency of image features. Therefore, the efficiency of image retrieval can be improved.

[0128] According to the method described in the above embodiments, the following will be further described in detail by way of examples.

[0129] In this embodiment, the image retrieval method can be executed by an electronic device, and the electronic device can be described as a server.

[0130] (1) The server trains an image retrieval model

[0131] S1. The server obtains an image sample set.

[0132] For example, the server can obtain original image samples, collect images misretrieved by the second image retrieval network in a preset image retrieval model, obtain abnormal image sample pairs, where the second image retrieval network is trained from the original image sample pairs, obtain generalized image samples of the original image sample pairs, and obtain generalized image sample pairs, and use the abnormal image sample pairs and the generalized image sample pairs as new target image sample pairs.

[0133] S2. The server extracts image sample triples for training from the original image sample pairs and the target image sample pairs.

[0134] For example, the server can batch the image sample set, with every bs image samples in a batch. In each batch (bs) of samples, an image sample is randomly selected as the target image sample x, and the distance between it and the remaining (bs - 1) image samples (one image sample is randomly selected from each pair) in the image sample pairs is calculated. The samples are sorted in ascending order of distance, and the TOP N samples are taken as negative image samples, and are respectively combined with the positive image samples in x to form triples, so as to extract the image sample triples corresponding to the image sample set.

[0135] S3. The server trains the first image retrieval network of the preset image retrieval model based on the original image sample pairs and the image sample triples to obtain the trained first image retrieval network.

[0136] For example, taking the first image retrieval network as a residual network (resnet101), the process of the server training the first image retrieval network can be as follows:

[0137] (1) Parameter initialization: Conv1 - Conv5 are initialized with the parameters of the second image retrieval network of the preset image retrieval model.

[0138] (2) Set learning parameters: Set the learning parameters as shown in Table 1 and Table 2, that is, all network parameters need to be learned.

[0139] (3) Learning rate: The basic feature and the image feature layer both adopt a learning rate of lr1 = 0.0005.

[0140] (4) Learning process: For the full amount of data, perform epoch (one round) iteration; process the full amount of samples once per round of iteration.

[0141] (5) The specific operations in each round of iteration are as follows: Divide every batch - size samples of the full amount of original image sample pairs into Nb batches. For each batch, perform the following training process:

[0142] A. Network forward: During training, the first image retrieval network performs forward calculation on an input original image sample to obtain a prediction result and outputs the original image feature.

[0143] B. Loss calculation: Find the original image sample triples in the batch samples and calculate the original triple loss information of the target original image features of the original image sample triples, which can be specifically shown in formula (1).

[0144] C. Update of the network parameters of the first image retrieval network: Adopt the SGD stochastic gradient descent algorithm or other gradient descent algorithms to perform gradient backward calculation on the original triple loss information to obtain the updated values of all network parameters and update the first retrieval network.

[0145] D. Stopping condition (convergence condition): Record the average loss information of the first image retrieval network in each epoch. When the average loss information of the first image retrieval network has not decreased for 5 consecutive rounds (epoch), stop the training of the first image retrieval network to obtain the trained first image retrieval network.

[0146] S4. The server jointly trains the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples to obtain the trained image retrieval model.

[0147] For example, taking the original image sample pair as the old image sample, the target image sample pair as the new image sample, the first image retrieval network as the new network, and the second image retrieval network as the old network, the server input samples pass through both the old and new networks simultaneously, generating old and new image features (embeddings). The old image features pass through the mapping module to generate mapped image features (second image features). Using the target triple loss information of the first image features, the mapping structure, and the mixed triple loss information of the new features as the supervision information, the feature mapping network and the new network (the trained first image retrieval network) are updated. The specific training process can be as follows:

[0148] 1) Parameter initialization: Use the learning result of the first new model (i.e., the trained first image retrieval network) as the initialization weight of the current training model.

[0149] 2) Set learning parameters: Set the learning parameters as shown in Table 1-2-3, that is, all network parameters need to be learned.

[0150] 3) Learning rate: For the basic features and the embedding layer, the learning rate of lr1 = 0.0005 is used.

[0151] 4) Learning process: Perform epoch rounds of iteration on the full amount of data; process the full amount of samples once per round of iteration;

[0152] 5) The specific operations in each round of iteration are as follows: Divide the full amount of samples (old and new image sample pairs) into batches of every batch-size samples, and divide them into Nb batches. For each batch, perform the following:

[0153] (1) Model forward: During training, the neural network performs forward calculation on an input image to obtain the prediction result and outputs the embedding result (the first image feature and the second image feature).

[0154] (2) Loss calculation: Calculate the triples according to the above method in the batch samples. According to the composition of the triple samples (the combination of badcase flags of the three images in the triple), calculate the loss of a certain triple, and sum up all the triple loss information to obtain the target loss information.

[0155] (3) Model parameter update: Use the SGD stochastic gradient descent method (or other convergence algorithms) to perform backward gradient calculation on the target loss information in the previous step to obtain the updated values of all model parameters, and update the feature mapping network and the trained first image retrieval network.

[0156] 6) Stopping condition: Record the average loss information of the model in each epoch. When the average loss information of the model does not decrease for 5 consecutive epochs or the preset number of epochs, stop the joint training of the first image retrieval network and the feature mapping network, and thus obtain the trained image retrieval model, which includes the updated first image retrieval network and the trained feature mapping network.

[0157] Among them, there are various ways to calculate the target loss information when jointly training the trained first image retrieval network and the feature mapping network. For example, the server determines the triplet loss information of the target image sample triplet based on the first image feature to obtain the target triplet loss information. The server selects any one image sample from the mixed image sample triplet as the current image sample for calculating the triplet loss information, calculates the triplet loss information of the mixed image sample triplet based on the current image sample and the second image feature to obtain the initial mixed triplet loss information, and returns to execute the step of selecting any one image sample from the mixed image sample triplet as the current image sample for calculating the triplet loss information until all image samples in the mixed image sample triplet are used as the current image sample to calculate the triplet loss information, so as to obtain the initial mixed triplet loss information corresponding to each image sample in the mixed image sample triplet, and fuse the initial mixed triplet loss information to obtain the mixed triplet loss information.

[0158] The server obtains the weighting coefficients of the target triplet loss information and the mixed triplet loss information, weights the target triplet loss information and the mixed triplet loss information respectively according to the weighting coefficients, and fuses the weighted target triplet loss information and the mixed triplet loss information to obtain the target loss information corresponding to the image sample set, which can be specifically shown in formula (2). Considering that the update of the new feature (the first image feature) needs to give priority to ensuring the retrieval effect of the new feature, a small weight is used to weight the mapped feature. According to the above principle of the mixed triplet loss information, the number of samples of the mixed triplet loss information is three times that of the target triplet loss information, and the calculation formula of the target loss information can also be shown in formula (3).

[0159] (2) The server uses the trained image retrieval model for image retrieval

[0160] Among them, the trained image retrieval model can include the updated first image retrieval network and the trained feature mapping network.

[0161] As Figure 6 shown, an image retrieval method has the following specific process:

[0162] 201. The server obtains the original image, the original image features corresponding to the original image, and at least one new target image.

[0163] For example, for the original image and the original image features, the server can directly obtain the original image and the original image features, or alternatively, it can also obtain the original image, extract features from the original image using the original image retrieval model, and obtain the original image features corresponding to the original image.

[0164] For the target image, the server can extract features from the original image using the original image retrieval model to obtain the original image features corresponding to the original image, input the target image to be retrieved into the original image retrieval model, and the original image retrieval model retrieves candidate images in the original image based on the original image features. When the retrieved candidate image is abnormal or incorrect, the target image to be retrieved is used as the target image that needs to be added. Or, it can obtain the image type of the original image, and based on this image type, screen out images different from this image type in the preset image database as the generalization image of the original image, and use this generalization image as the target image. Or, it can also receive the image that the user uploads through the terminal and needs to be added to the original image, thereby obtaining the target image.

[0165] 202. The server extracts features from the target image to obtain the target image features corresponding to the target image.

[0166] For example, the server can use the first image retrieval network of the trained image retrieval model to extract features from the target image, thereby obtaining the target image features corresponding to the target image.

[0167] 203. The server updates the original image features according to the target image features to obtain the updated image features.

[0168] For example, the server can obtain the original feature space information of the original image features and the target feature space information of the target image features, compare the original feature space information and the target feature space information, determine the mapping parameters for mapping the original image features based on the comparison result, and based on this mapping parameter, use the feature mapping network to map the original image features to the feature space of the target image features, thereby obtaining the updated image features. Or, based on the comparison result, screen out the mapping parameters corresponding to the comparison result in the mapping parameter set of the feature mapping network to obtain the target mapping parameter, and based on this target mapping parameter, use the feature mapping network to map the original image features to the feature space of the target image features, thereby obtaining the updated image features.

[0169] 204. The server fuses the target image features and the updated image features to obtain the full amount of image features for image retrieval.

[0170] For example, the server updates a preset image feature set based on the updated image features, obtains an updated image feature set, and adds the target image features to the updated image feature set to obtain a full set of image features. Alternatively, the server can also directly merge the target image features and the updated image features to obtain a full set of image features.

[0171] 205. When receiving an image retrieval request, the server extracts features from the image to be retrieved carried in the image retrieval request.

[0172] For example, the server can use the first image retrieval network in the trained image retrieval model to extract features from the image to be retrieved, obtaining the image features to be retrieved of the image to be retrieved. Here, the first image retrieval network in the trained image retrieval model can be the first image retrieval network obtained by jointly training the trained first image retrieval network.

[0173] 206. The server retrieves the target retrieval image from the original image and the target image according to the extracted image features to be retrieved and the full set of image features.

[0174] For example, the server can calculate the feature distance between the image features to be retrieved and each image feature in the full set of image features to obtain a first feature distance, sort the first feature distance, and retrieve the target retrieval image from the original image and the target image according to the sorting result. Alternatively, the server can also update the cluster center corresponding to the original image based on the full set of image features, filter out the target index in the index corresponding to the updated distance center based on the feature distance between the updated cluster center and the image features to be retrieved, and retrieve the target retrieval image from the original image and the target image according to the target index.

[0175] Optionally, after retrieving the target retrieval image corresponding to the image to be retrieved, the server can also classify or identify the image to be retrieved according to the category of the target retrieval image. Here, the classification or identification mainly refers to the identification at the category level of the image to be retrieved, without considering the specific instance of the object, but only considering the category of the object (such as people, dogs, cats, birds, etc.) for identification and giving the category to which the object belongs. There can be multiple ways to perform the classification or identification. For example, when the number of target retrieval images is 1, the classification or identification result of the image to be retrieved at this time can be the category of the target retrieval image. When the number of target retrieval images is multiple, the target retrieval images can be classified, and then, the number of images of each category of target retrieval images is obtained. The target image category with the largest number of images is selected from the image categories as the classification or identification result of the image to be retrieved. Or, when the number of target retrieval images is multiple, the target retrieval images are classified, and then, the sum of the feature distances corresponding to each image category is calculated, and the image category with the smallest sum of feature distances is used as the classification or identification result of the image to be retrieved.

[0176] As can be seen from the above, in this embodiment, after the server obtains the original image, the original image features corresponding to the original image, and at least one new target image, it extracts the features of the target image to obtain the target image features corresponding to the target image. Then, it updates the original image features according to the target image features to obtain the updated image features. Then, it fuses the target image features and the updated image features to obtain the full-scale image features for image retrieval. When receiving an image retrieval request, it extracts the features of the image to be retrieved carried in the image retrieval request, and retrieves the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features. Since this solution can migrate the original image features to the feature space of the target image features, it can quickly update the image features for image retrieval without having to re-extract the features of the original image, avoiding waste of computing resources, and can also improve the extraction efficiency of image features. Therefore, it can improve the efficiency of image retrieval.

[0177] To better implement the above method, an embodiment of the present invention further provides an image retrieval device, which can be integrated in a network device, such as a server or a terminal device.

[0178] For example, as Figure 7 shown, the image retrieval device may include an acquisition unit 301, an extraction unit 302, an update unit 303, a fusion unit 304, and a retrieval unit 305, as follows:

[0179] (1) Acquisition unit 301;

[0180] An acquisition unit 301, configured to acquire an original image, original image features corresponding to the original image, and at least one newly added target image, where the target image is an image newly added due to an incorrect retrieval of the original image.

[0181] For example, the acquisition unit 301 may specifically be configured to acquire an original image, extract features of the original image using an original image retrieval model to obtain original image features corresponding to the original image, input a target image to be retrieved into the original image retrieval model, and retrieve a candidate image from the original image by the original image retrieval model according to the original image features. When the retrieved candidate image is abnormal or incorrect, the target image to be retrieved is used as the target image to be newly added. Alternatively, the acquisition unit 301 may acquire the image type of the original image, and according to the image type, screen out an image different from the image type in a preset image database as a generalization image of the original image, and use the generalization image as the target image. Alternatively, the acquisition unit 301 may also receive an image uploaded by a user through a terminal and to be newly added to the original image, so as to obtain the target image.

[0182] (2) An extraction unit 302;

[0183] The extraction unit 302 is configured to extract features of the target image to obtain target image features corresponding to the target image.

[0184] For example, the extraction unit 302 may specifically be configured to extract features of the target image using a first image retrieval network of a trained image retrieval model, so as to obtain target image features corresponding to the target image.

[0185] (3) An update unit 303;

[0186] The update unit 303 is configured to update the original image features according to the target image features to obtain updated image features, where the updated image features and the target image features are image features in the same feature space.

[0187] For example, the update unit 303 may specifically be configured to obtain original feature space information of the original image features and target feature space information of the target image features, compare the original feature space information and the target feature space information, and based on the comparison result, map the original image features to the feature space of the target image features using a feature mapping network of a trained image retrieval model to obtain updated image features.

[0188] (4) A fusion unit 304;

[0189] The fusion unit 304 is configured to fuse the target image features and the updated image features to obtain all-image features for image retrieval.

[0190] For example, the fusion unit 304 can be specifically used to update a preset image feature set based on the updated image features to obtain an updated image feature set, add the target image features to the updated image feature set to obtain the full-scale image features, or directly merge the target image features and the updated image features to obtain the full-scale image features.

[0191] (5) Retrieval unit 305;

[0192] The retrieval unit 305 is used to extract features of the image to be retrieved carried in the image retrieval request when receiving the image retrieval request, and retrieve the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features.

[0193] For example, the retrieval unit 305 can be specifically used to extract features of the image to be retrieved by using the first image retrieval network in the trained image retrieval model when receiving the image retrieval request, to obtain the features of the image to be retrieved of the image to be retrieved. Calculate the feature distance between the features of the image to be retrieved and each image feature in the full-scale image features to obtain the first feature distance, sort the first feature distance, and retrieve the target retrieval image from the original image and the target image according to the sorting result, or update the clustering center corresponding to the original image according to the full-scale image features, and based on the feature distance between the updated clustering center and the features of the image to be retrieved, filter out the target index in the index corresponding to the updated distance center, and retrieve the target retrieval image from the original image and the target image according to the target index.

[0194] Optionally, the image retrieval device may further include a training unit 306, as Figure 8 shown, specifically as follows:

[0195] The training unit 306 is used to train the image retrieval model to obtain a trained image retrieval model.

[0196] For example, the training unit 306 can be specifically used to obtain an image sample set, which includes original image sample pairs and newly added target image sample pairs, extract image sample triples for training from the original image sample pairs and the target image sample pairs, train the first image retrieval network of the preset image retrieval model according to the original image sample pairs and the image sample triples to obtain a trained first image retrieval network, and jointly train the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples to obtain a trained image retrieval model.

[0197] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments and will not be elaborated herein.

[0198] As can be seen from the above, in this embodiment, after obtaining the original image, the original image features corresponding to the original image, and at least one newly added target image, feature extraction is performed on the target image to obtain the target image features corresponding to the target image. Then, the original image features are updated according to the target image features to obtain the updated image features. Then, the target image features and the updated image features are fused to obtain the full-scale image features for image retrieval. When an image retrieval request is received, feature extraction is performed on the image to be retrieved carried in the image retrieval request, and according to the extracted features of the image to be retrieved and the full-scale image features, the target retrieval image is retrieved from the original image and the target image. Since this solution can migrate the original image features to the feature space of the target image features, the image features for image retrieval can be updated quickly without re-extracting the features of the original image, avoiding waste of computing resources, and moreover, the extraction efficiency of the image features can be improved. Therefore, the efficiency of image retrieval can be improved.

[0199] An embodiment of the present invention further provides an electronic device, as Figure 9 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:

[0200] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 9 the structural diagram of the electronic device shown in

[0201] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Among them:

[0202] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.

[0203] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0204] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0205] Although not shown, in some embodiments, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0206] Obtain the original image, the original image features corresponding to the original image, and at least one newly added target image. The target image is an image newly added due to incorrect retrieval of the original image. Extract features from the target image to obtain the target image features corresponding to the target image. Update the original image features according to the target image features to obtain updated image features. The updated image features and the target image features are image features in the same feature space. Fuse the target image features and the updated image features to obtain the full amount of image features for image retrieval. When receiving an image retrieval request, extract features from the image to be retrieved carried in the image retrieval request, and retrieve the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full amount of image features.

[0207] For example, the electronic device obtains the original image, extracts features from the original image using the original image retrieval model to obtain the original image features corresponding to the original image, inputs the target image to be retrieved into the original image retrieval model, and the original image retrieval model retrieves candidate images from the original image according to the original image features. When the retrieved candidate images are abnormal or incorrect, the target image to be retrieved is used as the target image that needs to be newly added. Or, obtain the image type of the original image, and according to the image type, screen out images different from the image type in the preset image database as the generalization images of the original image, and use the generalization images as the target images. Or, it is also possible to receive the images that the user uploads through the terminal and needs to be newly added to the original image, so as to obtain the target images. Use the first image retrieval network of the trained image retrieval model to extract features from the target image, so as to obtain the target image features corresponding to the target image. Obtain the original feature space information of the original image features and the target feature space information of the target image features, and compare the original feature space information and the target feature space information. Based on the comparison result, use the feature mapping network of the trained image retrieval model to map the original image features to the feature space of the target image features to obtain updated image features. Based on the updated image features, update the preset image feature set to obtain an updated image feature set, and add the target image features to the updated image feature set to obtain the full amount of image features. Or, it is also possible to directly merge the target image features and the updated image features to obtain the full amount of image features. When receiving an image retrieval request, extract features from the image to be retrieved carried in the image retrieval request, and retrieve the target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full amount of image features.

[0208] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.

[0209] As can be seen from the above, in the embodiment of the present invention, after obtaining the original image, the original image features corresponding to the original image, and at least one newly added target image, feature extraction is performed on the target image to obtain the target image features corresponding to the target image. Then, the original image features are updated according to the target image features to obtain the updated image features. Then, the target image features and the updated image features are fused to obtain the full amount of image features for image retrieval. When an image retrieval request is received, feature extraction is performed on the image to be retrieved carried in the image retrieval request, and according to the extracted image features to be retrieved and the full amount of image features, the target retrieval image is retrieved from the original image and the target image. Since this solution can transfer the original image features to the feature space of the target image features, the image features for image retrieval can be quickly updated without having to re-extract the features of the original image, avoiding waste of computing resources, and also improving the extraction efficiency of the image features. Therefore, the efficiency of image retrieval can be improved.

[0210] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0211] Therefore, the embodiment of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the image retrieval methods provided by the embodiment of the present invention. For example, the instructions can execute the following steps:

[0212] Obtain the original image, the original image features corresponding to the original image, and at least one newly added target image. The target image is an image newly added due to an incorrect retrieval of the original image. Perform feature extraction on the target image to obtain the target image features corresponding to the target image. Update the original image features according to the target image features to obtain the updated image features. The updated image features and the target image features are image features in the same feature space. Fuse the target image features and the updated image features to obtain the full amount of image features for image retrieval. When an image retrieval request is received, perform feature extraction on the image to be retrieved carried in the image retrieval request, and according to the extracted image features to be retrieved and the full amount of image features, retrieve the target retrieval image from the original image and the target image.

[0213] For example, an electronic device obtains an original image, extracts features from the original image using an original image retrieval model to obtain original image features corresponding to the original image, inputs a target image to be retrieved into the original image retrieval model, and the original image retrieval model retrieves a candidate image from the original image based on the original image features. When the retrieved candidate image is abnormal or incorrect, the target image to be retrieved is used as the target image to be newly added. Alternatively, the image type of the original image is obtained, and according to this image type, an image different from this image type is selected from a preset image database as the generalization image of the original image, and this generalization image is used as the target image. Alternatively, an image that needs to be newly added to the original image uploaded by the user through a terminal can also be received to obtain the target image. The first image retrieval network of the trained image retrieval model is used to extract features from the target image to obtain target image features corresponding to the target image. The original feature space information of the original image features and the target feature space information of the target image features are obtained, and the original feature space information and the target feature space information are compared. Based on the comparison result, the feature mapping network of the trained image retrieval model maps the original image features to the feature space of the target image features to obtain updated image features. Based on the updated image features, a preset image feature set is updated to obtain an updated image feature set, and the target image features are added to the updated image feature set to obtain full-scale image features. Alternatively, the target image features and the updated image features can also be directly merged to obtain full-scale image features. When an image retrieval request is received, features are extracted from the image to be retrieved carried in the image retrieval request, and based on the extracted image features to be retrieved and the full-scale image features, a target retrieved image is retrieved from the original image and the target image.

[0214] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.

[0215] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0216] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the image retrieval methods provided by the embodiments of the present invention, the beneficial effects that can be achieved by any of the image retrieval methods provided by the embodiments of the present invention can be realized. For details, reference can be made to the previous embodiments and will not be elaborated here.

[0217] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various alternative implementations of the above-mentioned image retrieval aspect, image recognition aspect, or image classification.

[0218] The above has introduced in detail an image retrieval method, apparatus, and computer-readable storage medium provided by an embodiment of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An image retrieval method, characterized in that, it includes: obtaining an original image, the original image features corresponding to the original image, and at least one newly added target image, where the target image is an image newly added due to an incorrect retrieval of the original image; performing feature extraction on the target image to obtain target image features corresponding to the target image; updating the original image features according to the target image features to obtain updated image features, including: mapping the original image features to the feature space of the target image features to obtain the updated image features; the updated image features and the target image features are image features in the same feature space; fusing the target image features and the updated image features to obtain full-scale image features for image retrieval; when receiving an image retrieval request, performing feature extraction on the image to be retrieved carried in the image retrieval request, and retrieving a target retrieval image from the original image and the target image according to the extracted features of the image to be retrieved and the full-scale image features.

2. The image retrieval method according to claim 1, characterized in that, the updating the original image features according to the target image features to obtain updated image features includes: obtaining the original feature space information of the original image features and the target feature space information of the target image features, and comparing the original feature space information and the target feature space information; based on the comparison result, using the feature mapping network of the trained image retrieval model to map the original image features to the feature space of the target image features to obtain updated image features.

3. The image retrieval method according to claim 2, characterized in that, before the using the feature mapping network of the trained image retrieval model to map the original image features to the feature space of the target image features to obtain updated image features based on the comparison result, it further includes: obtaining an image sample set, where the image sample set includes original image sample pairs and newly added target image sample pairs; extracting image sample triples for training from the original image sample pairs and the target image sample pairs; training a first image retrieval network of a preset image retrieval model according to the original image sample pairs and the image sample triples to obtain a trained first image retrieval network; jointly training the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples to obtain a trained image retrieval model.

4. The image retrieval method according to claim 3, characterized in that, the obtaining an image sample set includes: obtaining original image sample pairs, and collecting images retrieved incorrectly by a second image retrieval network in the preset image retrieval model to obtain abnormal image sample pairs, where the second image retrieval network is trained from the original image sample pairs. Obtain the generalized images of the original image sample pairs to get generalized image sample pairs, and use the abnormal image sample pairs and the generalized image sample pairs as the newly added target image sample pairs; Fuse the target image sample pairs and the original image sample pairs to obtain the image sample set.

5. The image retrieval method according to claim 3, characterized in that, The extracting of the image sample triples for training from the original image sample pairs and the target image sample pairs includes: Screen out the target image samples from the original image sample pairs and the target image sample pairs; Calculate the image distances between the target image samples and the remaining image samples in the image sample set respectively; Based on the image distances, screen out a preset number of image negative samples corresponding to the target image samples in the image sample set, and add the image negative samples to the image sample pairs corresponding to the target image samples to obtain the image sample triples of the target image samples.

6. The image retrieval method according to claim 3, characterized in that, The training of the first image retrieval network of the preset image retrieval model according to the original image sample pairs and the image sample triples to obtain the trained first image retrieval network includes: Use the first image retrieval network of the preset image retrieval model to extract features from the original image samples in the original image sample pairs to obtain the original image sample features; Screen out the target original image features corresponding to the image sample triples from the original image sample features; Calculate the feature distances between the target original image features to obtain the original triple loss information corresponding to the image sample triples; Converge the first image retrieval network based on the original triple loss information to obtain the trained first image retrieval network.

7. The image retrieval method according to claim 4, characterized in that, The joint training of the trained first image retrieval network and the feature mapping network in the preset image retrieval model based on the image sample set and the image sample triples to obtain the trained image retrieval model includes: Use the trained first image retrieval network to extract features from the image samples in the image sample set to obtain the first image features; Use the second image retrieval network to extract features from the image samples in the image sample set, and use the feature mapping network in the preset image retrieval model to map the extracted features to obtain the second image features; Converge the trained first image retrieval network and the feature mapping network according to the first image features, the second image features and the image sample triples to obtain the trained image retrieval model.

8. The image retrieval method according to claim 7, characterized in that, The converging of the trained first image retrieval network and the feature mapping network according to the first image features, the second image features and the image sample triples to obtain the trained image retrieval model includes: Select the triplets composed of the image samples in the target image sample pair from the image sample triplets to obtain target image sample triplets; Select the triplets composed of the image samples in the target image sample pair and the image samples in the original image sample pair from the image sample triplets to obtain mixed image sample triplets; Determine the target loss information corresponding to the image sample set according to the target image sample triplets, the mixed image sample triplets, the first image feature, and the second image feature; Converge the trained first image retrieval network and the feature mapping network according to the target loss information to obtain a trained image retrieval model.

9. The image retrieval method according to claim 8, wherein, the determining the target loss information corresponding to the image sample set according to the target image sample triplets, the mixed image sample triplets, the first image feature, and the second image feature includes: Determine the triplet loss information of the target image sample triplets according to the first image feature to obtain target triplet loss information; Based on the second image feature, determine the triplet loss information of the mixed image sample triplets to obtain mixed triplet loss information; Fuse the target triplet loss information and the mixed triplet loss information to obtain the target loss information corresponding to the image sample set.

10. The image retrieval method according to claim 9, wherein, the determining the triplet loss information of the mixed image sample triplets based on the second image feature to obtain mixed triplet loss information includes: Select any one image sample from the mixed image sample triplets as the current image sample for calculating the triplet loss information; Calculate the triplet loss information of the mixed image sample triplets according to the current image sample and the second image feature to obtain initial mixed triplet loss information; Return to execute the step of selecting any one image sample from the mixed image sample triplets as the current image sample for calculating the triplet loss information until all the image samples in the mixed image sample triplets are used as the current image sample to calculate the triplet loss information, so as to obtain the initial mixed triplet loss information corresponding to each image sample in the mixed image sample triplets; Fuse the initial mixed triplet loss information to obtain mixed triplet loss information.

11. The image retrieval method according to any one of claims 1 to 10, wherein, the retrieving the target retrieval image from the original image and the target image according to the extracted image feature to be retrieved and the full amount of image features includes: Calculate the feature distance between the image feature to be retrieved and each image feature in the full amount of image features to obtain a first feature distance; Sort the first feature distances, and retrieve the target retrieval image from the original image and the target image according to the sorting result.

12. An image retrieval device, wherein, comprising: An acquisition unit, configured to acquire an original image, original image features corresponding to the original image, and at least one newly added target image, where the target image is an image newly added due to an incorrect retrieval of the original image; An extraction unit, configured to perform feature extraction on the target image to obtain target image features corresponding to the target image; An update unit, configured to update the original image features according to the target image features to obtain updated image features, including: mapping the original image features to the feature space of the target image features to obtain the updated image features; the updated image features and the target image features are image features in the same feature space; A fusion unit, configured to fuse the target image features and the updated image features to obtain full-scale image features for image retrieval; A retrieval unit, configured to, when receiving an image retrieval request, perform feature extraction on the image to be retrieved carried in the image retrieval request, and retrieve a target retrieval image from the original image and the target image according to the extracted image features to be retrieved and the full-scale image features.

13. An electronic device, characterized in that, it includes a processor and a memory, the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps in the image retrieval method according to any one of claims 1 to 11.

14. A computer program product, including a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, the steps in the image retrieval method according to any one of claims 1 to 11 are implemented.

15. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the image retrieval method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Object category identification method and device and server

    CN112733969A