Image retrieval method, device, storage medium and electronic device
By training the target model and combining the similarity comparison between image and user information features, the problem of insufficient search accuracy caused by ignoring the style of user's works in the prior art is solved, and accurate image retrieval in content and style is achieved.
Patent Information
- Application Number
- CN202210068511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-20
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-01-20
AI Technical Summary
When searching in massive user-made original content images, the prior art only considers the similarity of the image content and ignores the style of the user's work, resulting in insufficient retrieval accuracy, which is difficult to apply especially when the image perspective and content display morphology changes strongly.
By training the target model, the user content similarity of the positive and negative sample images is used, combined with the image and user information characteristics, the feature representation of the image to be retrieved is obtained, and the similarity is compared with the user image set to determine the target image.
The accuracy of image retrieval is improved, so that the retrieved images are not only similar in content, but also similar to the image to be retrieved in the user's style, which improves the accuracy of the search.
Smart Images

Figure CN114428872B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image retrieval method, device, storage medium, and electronic device. Background Art
[0002] Currently, when retrieving data with perceptually similar content to a specified image from massive user-generated content (UGC) images, only the similarity in image content is considered, without understanding the user's work style, making it impossible to guarantee the accuracy of the retrieval. Summary of the Invention
[0003] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides an image retrieval method, comprising:
[0005] Inputting an image to be retrieved into a pre-trained target model, and obtaining a feature representation of the image to be retrieved output by the target model, wherein the training samples of the target model include positive sample images and negative sample images, wherein the positive sample images and the negative sample images are determined from a plurality of sample images based on similarities between user content to which the plurality of sample images belong, wherein the user content includes sample images and user information corresponding to the sample images;
[0006] performing a similarity comparison between the feature expression and the feature expression of each image in a user image set, respectively, to obtain a plurality of similarity results, wherein the user image set includes a plurality of images generated by a plurality of users;
[0007] A target image is determined from the user image set according to the multiple similarity results.
[0008] In a second aspect, the present disclosure provides an image retrieval device, comprising:
[0009] a feature identification acquisition module, configured to input an image to be retrieved into a pre-trained target model and obtain a feature representation of the image to be retrieved output by the target model, wherein the training samples of the target model include positive sample images and negative sample images, wherein the positive sample images and the negative sample images are determined from a plurality of sample images based on similarities between user content to which the sample images belong, wherein the user content includes sample images and user information corresponding to the sample images;
[0010] a similarity result acquisition module, configured to perform a similarity comparison between the feature expression and the feature expression of each image in a user image set, to obtain a plurality of similarity results, wherein the user image set includes a plurality of images generated by a plurality of users;
[0011] The target image determination module is configured to determine a target image from the target image set according to the multiple similarity results.
[0012] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.
[0013] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0014] a storage device having a computer program stored thereon;
[0015] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect.
[0016] The image retrieval method, apparatus, storage medium, and electronic device provided herein input an image to be retrieved into a pre-trained target model to obtain a feature representation of the image to be retrieved output by the target model. The training samples of the target model include positive sample images and negative sample images. The positive sample images and negative sample images are determined from multiple sample images based on the similarity between the user content to which the sample images belong. The user content includes the sample images and user information corresponding to the sample images. The feature representation is then compared with the feature representation of each image in a user image collection to obtain multiple similarity results. The user image collection includes multiple images generated by multiple users. Finally, based on the multiple similarity results, a target image is determined from the user image collection. Because the target model is trained based on sample images and user information in the user content, the feature representation output by the target model can also reflect the image from two dimensions: the feature dimension of the user information and the feature dimension of the image sample. Therefore, when performing image retrieval using the feature representation output by the target model, the retrieved image can be similar not only in content to the image to be retrieved, but also in style to the image to be detected by the user, thereby improving retrieval accuracy.
[0017] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:
[0019] Figure 1 The figure is a flowchart of an image retrieval method according to an exemplary embodiment.
[0020] Figure 2 The figure is a flowchart of an image retrieval method according to another exemplary embodiment.
[0021] Figure 3 is based on Figure 2 A network connection diagram between multiple sample images shown in the embodiment.
[0022] Figure 4 is based on Figure 2 The embodiment shows a flow chart of step 250 of the image retrieval method.
[0023] Figure 5 The figure is a block diagram of an image retrieval device according to an exemplary embodiment.
[0024] Figure 6 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0025] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0026] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0027] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] Deep learning has developed rapidly in recent years, attracting widespread attention both domestically and internationally. With the continuous advancement of deep learning technology and the continuous improvement of data processing capabilities, more and more deep learning algorithms are being applied to image processing and computer vision. Image retrieval is an important area of computer vision.
[0032] Among them, it is often necessary to find data with similar sensory perception to the specified image content in massive user-generated content images, such as finding the same shooting scene as the user-uploaded photos, or similar subjects.
[0033] In related technologies, image learning methods based on comparative learning can learn a certain degree of discriminative features. However, due to the limitations of the implementation categories of data augmentation, the definition of similar images has certain limitations. It mainly focuses on the image content and lacks an abstract understanding of the user's work style. As a result, the retrieved similar images are only similar in content and cannot form a category in terms of user style. It is difficult to apply when the image perspective and content presentation form change drastically.
[0034] In addition, in related technologies, when training models based on deep learning methods, a large amount of data is required, and whether the image content is similar must be labeled, resulting in low training efficiency.
[0035] In response to the above problems, the present disclosure provides an image retrieval method, device, electronic device and storage medium, which can effectively improve the retrieval accuracy of user-generated images.
[0036] Figure 1 is a flow chart of an image retrieval method according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps:
[0037] 110. Input the image to be retrieved into a pre-trained target model to obtain the feature representation of the image to be retrieved output by the target model. The training samples of the target model include positive sample images and negative sample images. The positive sample images and negative sample images are determined from multiple sample images based on the similarity between the user content to which the multiple sample images belong. The user content includes the sample image and the user information corresponding to the sample image.
[0038] For example, the image retrieval method of this embodiment can be executed by an electronic device, a terminal device, an image processing device or apparatus, or other devices or apparatuses capable of executing this embodiment, without limitation. This embodiment is described with the electronic device as the executing device.
[0039] In some embodiments, a target model is configured in the electronic device. After the electronic device determines the image to be retrieved, the image to be retrieved can be input into the target model, and the target model can output a feature representation corresponding to the image to be retrieved (hereinafter referred to as the first feature representation), where the feature representation can be a specified number of floating-point arrays, such as 1200 floating-point arrays.
[0040] Among them, user information may include: user account, user name, etc., and may also be user portrait information that can represent the user's identity (such as age, gender, etc.).
[0041] The similarity between user contents can be determined from two aspects: the image features of the sample image and the information features of the user information. For example, for two user contents, the similarity between the image features and the similarity between the information features in the two user contents can be accumulated according to the preset weights to obtain the similarity.
[0042] The image to be retrieved may be an image that is custom-selected by the user from a plurality of existing images, or may be an image that is instantly generated by the electronic device based on some image attributes input by the user.
[0043] 120. Perform similarity comparison between the feature expression and the feature expression of each image in the user image set to obtain multiple similarity results. The user image set includes multiple images generated by multiple users.
[0044] In some embodiments, the user image set may be pre-stored in the electronic device, or the electronic device may obtain and update the user image set in real time from various platforms based on the user's user identity information (such as a user account). After obtaining the user image set, the electronic device may input each image in the user image set into the target model to obtain a feature representation (hereinafter referred to as a second feature representation) corresponding to each image in the user image set output by the target model. The electronic device may then perform a similarity comparison between the first feature representation and the plurality of second feature tables to obtain similarity results between the first feature representation and the plurality of second feature representations.
[0045] For example, the second feature representation includes second feature representation 1, second feature representation 2, second feature representation 3, and so on.
[0046] The electronic device can calculate the cosine similarity between the first feature representation and the second feature representation 1 to obtain similarity result 1, calculate the cosine similarity between the first feature representation and the second feature representation 2 to obtain similarity result 2, calculate the cosine similarity between the first feature representation and the second feature representation 3 to obtain similarity result 3, and so on, to obtain similarity results of the first feature representation corresponding to multiple second feature representations.
[0047] 130. Determine a target image from the user image set based on the multiple similarity results.
[0048] In some embodiments, the electronic device may select a second feature representation having the greatest similarity to the first feature representation from among multiple second feature representations based on multiple similarity results as the target feature representation, and determine the image corresponding to the target feature representation as the target image.
[0049] Continuing with the above example, if similarity result 2 has the greatest similarity among similarity result 1, similarity result 2, similarity result 3, etc., then the image corresponding to the second feature representation 2 can be determined as the target image.
[0050] As can be seen, in this embodiment, by inputting the image to be retrieved into a pre-trained target model, a feature representation of the image to be retrieved is obtained from the target model. The training samples of the target model include positive sample images and negative sample images. The positive and negative sample images are determined from multiple sample images based on the similarity between the user content to which the sample images belong. The user content includes the sample images and the user information corresponding to the sample images. The feature representation is then compared with the feature representation of each image in the user image collection to obtain multiple similarity results. The user image collection includes multiple images generated by multiple users. Finally, based on the multiple similarity results, the target image is determined from the user image collection. Because the target model is trained based on sample images and user information in the user content, the feature representation output by the target model can also reflect the image from two dimensions: the feature dimension of the user information and the feature dimension of the image sample. Therefore, when performing image retrieval using the feature representation output by the target model, the retrieved image can be similar not only in content to the image to be retrieved, but also in style to the image to be detected by the user, thereby improving retrieval accuracy.
[0051] Figure 2 is a flowchart of an image retrieval method according to another exemplary embodiment. Figure 2 As shown, the method may include the following steps:
[0052] 210. Obtain image features of each sample image among a plurality of sample images and information features of user information corresponding to each sample image.
[0053] In some embodiments, the specific implementation of step 210 includes:
[0054] A plurality of sample images are input into a pre-trained image feature extraction model, and the image features of each sample image output by the image feature extraction model are obtained, wherein the image feature extraction model is trained in a self-supervised manner.
[0055] For example, the electronic device can use a deep learning comparison learning method (such as MoCo) to train an image feature extraction model based on image data in a self-supervised manner. When using the image feature extraction model, the electronic device inputs multiple sample images into the image feature extraction model to extract the image features corresponding to each sample image.
[0056] In this embodiment, the image features of the sample image can be extracted quickly and effectively through the image feature extraction model. In addition, the image feature extraction model is trained based on a self-supervised method, which eliminates the dependence on data labeling, simplifies the training process, and improves training efficiency.
[0057] The image features may include but are not limited to: color features, texture features, shape features, etc.
[0058] In some embodiments, the user information may include a user ID number. Obtaining information features of the user information corresponding to each sample image in the plurality of sample images includes:
[0059] An information feature is generated based on the user ID, wherein the information feature may be a string of floating-point numbers corresponding to the user ID. For example, different user IDs may correspond to different floating-point numbers in advance, and the electronic device may determine the floating-point number corresponding to the user ID as the information feature of the user information.
[0060] 220. Perform similarity aggregation on the image features of each sample image to obtain a first aggregation result.
[0061] In some embodiments, the electronic device may aggregate sample images whose image features have similarities greater than a first specified similarity.
[0062] For example, the sample images include sample image 1, sample image 2, sample image 3, and sample image 4. Sample image 1 corresponds to image feature a1, sample image 2 corresponds to image feature a2, sample image 3 corresponds to image feature a3, and sample image 4 corresponds to image feature a4. Image feature a1 can be used as a cluster center. If the similarity between image features a2 and a3 and image feature a1 is equal to the first specified similarity, image features a1, a2, and a3 can be grouped as one class. Image feature a4 can be grouped as another class, thereby obtaining a first aggregation result.
[0063] Among them, the calculation formula of the cosine similarity between two image features can be as follows:
[0064] sim(f a ,f b )=||f a ∣∣·∣∣f b |||cos(θ);
[0065] Among them, f a is one of the image features, f b is another image feature.
[0066] 230. Perform similarity aggregation on the information features of the user information corresponding to each sample image to obtain a second aggregation result.
[0067] Continuing with the above example, let's assume that the sample images include sample image 1, sample image 2, sample image 3, and sample image 4. Sample image 1 corresponds to information feature b1, sample image 2 corresponds to image feature b2, sample image 3 corresponds to image feature b3, and sample image 4 corresponds to image feature b4. Image feature b1 can be used as a cluster center. If the similarity between image feature b2 and image feature a1 is greater than the second specified similarity, image features a1 and a2 can be grouped as one class. If the similarity between image features b3 and a4 is greater than the second specified similarity, image features b3 and a4 can be grouped as another class, thereby obtaining a second aggregation result.
[0068] 240. Determine the class to which the comprehensive features of each sample image belong based on the first aggregation result and the second aggregation result.
[0069] Among them, comprehensive features include image features and information features.
[0070] In one example, image features of the same user information may be clustered into one category, and then image features in the category with a similarity greater than a first specified similarity may be determined to be in the same category.
[0071] In some embodiments, the specific implementation of step 240 includes:
[0072] If the first sample image and the second sample image among the plurality of sample images belong to the same class in the first aggregation result and / or the second clustering result, it is determined that the comprehensive features of the first sample image and the comprehensive features of the second sample image belong to the same class.
[0073] Continuing with the above example, if sample image 1 and sample image 2 can be classified into the same class based on both image features and information features, then the comprehensive features of sample image 1 and sample image 2 can also be classified into the same class. Similarly, the class to which the comprehensive features of each sample image belong can be determined.
[0074] In another example, Figure 3 As shown in , a network connection diagram between multiple sample images can be constructed, where each sample image represents a node, and each two nodes can be connected by two edges, one edge represents the similarity of the image feature dimension, and the other edge represents the similarity of the information feature dimension. When the similarity of any edge does not meet the preset similarity condition, the edge can be deleted. According to the network connection diagram, the node with the largest connection can be regarded as a class, thereby obtaining multiple classes divided according to comprehensive features, such as Figure 3As shown, it includes the first category (sample image 1, sample image 2, sample image 3) and the second category (sample image 4). Similarly, if sample image 2 is also connected to sample image 5, then sample image 1 and sample image 5 are also in the same category.
[0075] 250. Determine positive sample images and negative sample images from the plurality of sample images according to the class to which the comprehensive feature of each sample image belongs and the preset benchmark comprehensive feature.
[0076] In some embodiments, as Figure 4 As shown, the specific implementation of step 250 may include:
[0077] 251. Determine, among the multiple sample images, a sample image corresponding to a comprehensive feature belonging to the same category as the preset reference comprehensive feature as a positive sample image.
[0078] 252. Determine, among the multiple sample images, a sample image corresponding to a comprehensive feature belonging to a different category from the preset reference comprehensive feature as a negative sample image.
[0079] In some embodiments, before step 250, the method further includes: obtaining a cluster center of each of the multiple classes to which all comprehensive features belong, and merging the classes corresponding to two cluster centers whose similarity is greater than a first similarity threshold among the multiple classes.
[0080] Among them, the cluster center in a class and other data in the class all meet the preset similarity requirements, and the other data in this class are called the associated data of the cluster center.
[0081] For example, category 1 includes cluster center a1 and associated data a1, and category 2 includes cluster center b1 and associated data b1. If the similarity between cluster center a1 and cluster center b1 is greater than a first similarity threshold, category 1 and category 2 can be merged.
[0082] In some embodiments, before obtaining the cluster center of each of the multiple classes to which all comprehensive features belong and merging the classes corresponding to two cluster centers whose similarity in the multiple classes is greater than a first similarity threshold, the method may further include:
[0083] It is determined that the number of multiple classes to which all comprehensive features belong is greater than or equal to a quantity threshold, so that similar classes can be merged when there are too many categories.
[0084] 260. The target model is obtained based on the training of positive sample images and negative sample images.
[0085] In some embodiments, the target model can use a common classification network as the basic structure, such as ResNet, DenseNet, etc.
[0086] For example, when training the target model, triplet loss can be used. The benchmark data a is selected, positive examples p are sampled from the data belonging to the same category as a, and negative examples n are sampled from the data not belonging to the same category as a. The loss is calculated as follows:
[0087] L=max(d(a,p)+margin,0);
[0088] Where L is the final loss, d(·) is the model output, a is the baseline data, p is the data of the same category as a, n is the data of a different category, and margin is a preset parameter representing the distance between positive and negative examples. When the final loss value L reaches a certain level, the target model training is complete.
[0089] For example, the benchmark data a can be randomly selected from the divided classes. For example, there are three classes: Q, W, and E. The benchmark data a can be randomly selected from the three classes, such as randomly selecting one from Q. The positive example is another one from Q, and the negative example is one from non-Q (W or E).
[0090] Optionally, when training the target model, a gradient descent training model can be used, and the Adam (Adaptive Moment Estimation) algorithm can be used as the optimization algorithm.
[0091] 270. Input the image to be retrieved into a pre-trained target model to obtain the feature representation of the image to be retrieved output by the target model.
[0092] Among them, the training samples of the target model include positive sample images and negative sample images. The positive sample images and negative sample images are determined from multiple sample images based on the similarity between the user content to which the multiple sample images belong. The user content includes the sample images and the user information corresponding to the sample images.
[0093] In some embodiments, before step 270, the method may further include:
[0094] Remove the loss calculation layer of the target model.
[0095] For example, during the evaluation phase, the electronic device may remove the loss calculation layer of the network and retain only the embedding output of the network as the feature representation of the image.
[0096] 280. Perform similarity comparison between the feature expression and the feature expression of each image in the user image set to obtain multiple similarity results. The user image set includes multiple images generated by multiple users.
[0097] 290. Determine a target image from the user image set based on the multiple similarity results.
[0098] In some embodiments, specific implementations of step 290 may include:
[0099] An image in the user image set whose similarity result is greater than or equal to a second similarity threshold is determined as a target image.
[0100] In other embodiments, specific implementations of step 290 may include:
[0101] The first N images with the largest similarity results in the user image set are determined as target images, where N is a positive integer.
[0102] Figure 5 is a block diagram of an image retrieval device according to an exemplary embodiment. Figure 5 As shown, the image retrieval device 300 may include: a feature identification acquisition module 310, a similarity result acquisition module 320, and a target image determination module 330, wherein:
[0103] The feature identification acquisition module 310 is used to input the image to be retrieved into a pre-trained target model and obtain the feature representation of the image to be retrieved output by the target model. The training samples of the target model include positive sample images and negative sample images. The positive sample images and negative sample images are determined from multiple sample images based on the similarity between the user content to which the multiple sample images belong. The user content includes the sample image and the user information corresponding to the sample image.
[0104] The similarity result acquisition module 320 is used to perform similarity comparison between the feature expression and the feature expression of each image in the user image set to obtain multiple similarity results. The user image set includes multiple images generated by multiple users.
[0105] The target image determination module 330 is configured to determine a target image from the target image set according to the multiple similarity results.
[0106] In some embodiments, the apparatus 300 further includes:
[0107] The feature extraction module is used to obtain the image features of each sample image in a plurality of sample images and the information features of the user information corresponding to each sample image.
[0108] The first aggregation module is used to perform similarity aggregation on the image features of each sample image to obtain a first aggregation result.
[0109] The second aggregation module is used to perform similarity aggregation on the information features of the user information corresponding to each sample image to obtain a second aggregation result.
[0110] The category determination module is used to determine the category to which the comprehensive features of each sample image belong based on the first aggregation result and the second aggregation result.
[0111] The positive and negative sample determination module is used to determine positive sample images and negative sample images from multiple sample images based on the class to which the comprehensive features of each sample image belong and the preset benchmark comprehensive features.
[0112] The model training module is used to train the target model based on positive sample images and negative sample images.
[0113] In some embodiments, the category determination module is specifically used to determine that the comprehensive features of the first sample image and the comprehensive features of the second sample image belong to the same category if the first sample image and the second sample image in the multiple sample images are of the same category in the first aggregation result and / or the second clustering result.
[0114] In some embodiments, the feature extraction module includes:
[0115] The image feature extraction submodule is used to input multiple sample images into a pre-trained image feature extraction model and obtain the image features of each sample image output by the image feature extraction surface model, wherein the image feature extraction model is trained in a self-supervised manner.
[0116] In some embodiments, the positive and negative sample determination module includes:
[0117] The positive sample determination submodule is used to determine, among multiple sample images, a sample image corresponding to a comprehensive feature belonging to the same category as the preset reference comprehensive feature as a positive sample image.
[0118] The negative sample determination submodule is used to determine, among the multiple sample images, sample images corresponding to comprehensive features belonging to different categories from the preset reference comprehensive features as negative sample images.
[0119] In some embodiments, the apparatus 300 further includes:
[0120] The cluster center acquisition module is used to obtain the cluster center of each class in the multiple classes to which all comprehensive features belong.
[0121] The merging module is used to merge the classes corresponding to two cluster centers whose similarity among the multiple classes is greater than a first similarity threshold.
[0122] In some embodiments, the target image determination module 330 includes:
[0123] The first determination submodule is configured to determine the first N images with the largest similarity results in the user image set as target images, where N is a positive integer.
[0124] In some embodiments, the target image determination module 330 includes:
[0125] The second determining submodule is configured to determine an image in the user image set whose similarity result is greater than or equal to a second similarity threshold as a target image.
[0126] In some embodiments, the apparatus 300 further includes:
[0127] The removal module is used to remove the loss calculation layer of the target model.
[0128] Reference below Figure 6 , which shows a schematic structural diagram of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0129] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0130] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0131] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0132] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0133] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0134] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0135] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: inputs the image to be retrieved into a pre-trained target model, obtains the feature representation of the image to be retrieved output by the target model, the training samples of the target model include positive sample images and negative sample images, the positive sample images and negative sample images are determined from multiple sample images based on the similarity between the user content to which the multiple sample images belong, and the user content includes the sample image and the user information corresponding to the sample image; compares the feature expression with the feature expression of each image in the user image set for similarity, and obtains multiple similarity results, the user image set includes multiple images generated by multiple users; and determines the target image from the user image set based on the multiple similarity results.
[0136] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0138] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. The name of a module does not, in some cases, limit the module itself. For example, a similarity result acquisition module may also be described as "a module that performs a similarity comparison between the feature expression and the feature expression of each image in a user image set, thereby obtaining multiple similarity results, where the user image set includes multiple images generated by multiple users."
[0139] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0140] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0142] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0143] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. An image retrieval method, characterized in that: include: Input the image to be retrieved into a pre-trained target model, and obtain the feature representation of the image to be retrieved output by the target model, wherein the training samples of the target model include positive sample images and negative sample images, wherein the positive sample images are sample images corresponding to comprehensive features belonging to the same category as the preset benchmark comprehensive features among multiple sample images, and the negative sample images are sample images corresponding to comprehensive features belonging to different categories than the preset benchmark comprehensive features among the multiple sample images, wherein the comprehensive features include image features of the sample images and information features of user information corresponding to the sample images; performing a similarity comparison between the feature representation and feature representations of each image in a user image set, respectively, to obtain a plurality of similarity results, wherein the user image set includes a plurality of images generated by a plurality of users; A target image is determined from the user image set according to the multiple similarity results.
2. The method according to claim 1, characterized in that Before inputting the image to be retrieved into the pre-trained target model and obtaining the feature representation of the image to be retrieved output by the target model, the method further includes: Obtaining an image feature of each sample image from a plurality of sample images and an information feature of user information corresponding to each sample image; Performing similarity aggregation on the image features of each sample image to obtain a first aggregation result; Performing similarity aggregation on the information features of the user information corresponding to each sample image to obtain a second aggregation result; Determining, based on the first aggregation result and the second aggregation result, the class to which the comprehensive feature of each sample image belongs; determining a positive sample image and a negative sample image from the plurality of sample images according to the class to which the comprehensive feature of each sample image belongs and a preset benchmark comprehensive feature; The target model is obtained by training based on the positive sample images and the negative sample images.
3. The method according to claim 2, characterized in that The determining, based on the first aggregation result and the second aggregation result, the class to which the comprehensive feature of each sample image belongs includes: If a first sample image and a second sample image among the plurality of sample images belong to the same category in the first aggregation result and / or the second aggregation result, it is determined that the comprehensive features of the first sample image and the comprehensive features of the second sample image belong to the same category.
4. The method according to claim 2, characterized in that The acquiring of the image feature of each sample image in the plurality of sample images includes: The plurality of sample images are input into a pre-trained image feature extraction model, and the image features of each sample image output by the image feature extraction model are obtained, wherein the image feature extraction model is trained in a self-supervised manner.
5. The method according to claim 2, characterized in that Before determining the positive sample image and the negative sample image from the plurality of sample images according to the class to which the comprehensive feature of each sample image belongs and the preset reference comprehensive feature, the method further includes: Get the cluster center of each class in the multiple classes to which all comprehensive features belong; The classes corresponding to two cluster centers whose similarity among the multiple classes is greater than a first similarity threshold are merged.
6. The method according to any one of claims 1 to 5, characterized in that Determining a target image from the user image set based on the multiple similarity results includes: The first N images with the largest similarity results in the user image set are determined as the target images, where N is a positive integer.
7. The method according to any one of claims 1 to 5, characterized in that Determining a target image from the user image set based on the multiple similarity results includes: An image in the user image set whose similarity result is greater than or equal to a second similarity threshold is determined as the target image.
8. The method according to any one of claims 1 to 5, characterized in that Before inputting the image to be retrieved into the pre-trained target model and obtaining the feature representation of the image to be retrieved output by the target model, the method further includes: Remove the loss calculation layer of the target model.
9. An image retrieval device, characterized in that: include: a feature identification acquisition module, configured to input an image to be retrieved into a pre-trained target model and obtain a feature representation of the image to be retrieved output by the target model, wherein the training samples of the target model include positive sample images and negative sample images, wherein the positive sample images are sample images corresponding to comprehensive features belonging to the same category as the preset reference comprehensive features among multiple sample images, and the negative sample images are sample images corresponding to comprehensive features belonging to different categories than the preset reference comprehensive features among the multiple sample images, wherein the comprehensive features include image features of the sample images and information features of user information corresponding to the sample images; a similarity result acquisition module, configured to perform a similarity comparison between the feature representation and the feature representation of each image in a user image set, to obtain a plurality of similarity results, wherein the user image set includes a plurality of images generated by a plurality of users; The target image determination module is configured to determine a target image from the target image set according to the multiple similarity results.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 8 are implemented.
11. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image retrieval method and system
CN108829848A
Image classification method and device, electronic equipment and storage medium
CN112434178A