Method for image search and image indexing, and electronic device

By using Semantic Alignment Network (SAN) to generate semantic embeddings based on semantic constraints, providing visual-semantic space, solving the problem of insufficient word space in existing text-to-photo retrieval methods and improving retrieval accuracy.

CN114945907BActive Publication Date: 2025-06-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180009096.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-12
Filing Date
2021-02-08
Publication Date
2025-06-27
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

Existing text-to-photo retrieval methods rely on encoding images into embedding vectors, but since word space is based on word co-occurrence information in the corpus, it is difficult to effectively capture the semantic similarity between words, resulting in the search results that do not meet expectations.

Method used

Semantic alignment network (SAN) is used to improve the shortcomings of word vectors, and generate semantic embeddings based on semantic constraints through SAN, thereby providing visual-semantic space for image search and indexing.

Benefits of technology

Improve the accuracy of text-to-photo retrieval, ensure the matching degree of query keywords with the corresponding image, and avoid inconsistent results due to insufficient word space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114945907B_ABST
    Figure CN114945907B_ABST
Patent Text Reader

Abstract

An image search method, comprising: obtaining a query keyword; and obtaining a target semantic embedding of the query keyword through a Semantic Alignment Network (SAN), and searching for at least one target image corresponding to the query keyword according to the target semantic embedding through the SAN. The SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space. The visual-semantic space defines a mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 975,565, filed on February 12, 2020, the content of which is hereby incorporated by reference in its entirety. Technical Field

[0003] This application relates to the technical field of image processing, and particularly to an image search method, an image indexing method, and an electronic device. Background Art

[0004] Due to the rapid increase in photos generated by mobile phone cameras, people's interest in text-to-photo retrieval has increased, and thus there is a need to effectively find the desired images from a large number of photos. Summary of the Invention

[0005] In a first aspect, this application proposes an image search method, including: obtaining a query keyword; and obtaining a target semantic embedding of the query keyword through a Semantic Alignment Network (SAN), and searching for at least one target image corresponding to the query keyword through the SAN according to the target semantic embedding; wherein, the SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space; the visual-semantic space defines a mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0006] In a second aspect, this application proposes an image indexing method, including: obtaining at least one image; and converting the at least one image into at least one image embedding through a Semantic Alignment Network (SAN), thereby providing a visual-semantic space; wherein, the visual-semantic space defines a mapping relationship between the at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0007] In a third aspect, this application proposes an electronic device, including a processor and a memory. The memory stores instructions, and when the instructions are executed by the processor, the processor executes the above method. Brief Description of the Drawings

[0008] In order to make the technical solutions described in the embodiments of this application clearer, the drawings used to describe the embodiments will be briefly described. Obviously, the described drawings are only for illustrating this application, rather than for limiting this application. It should be understood that those skilled in the art can obtain other drawings based on these drawings without creative efforts.

[0009] Figure 1It is a framework diagram of a Semantic Alignment Network (SAN) provided by an embodiment of the present application.

[0010] Figure 2 It is a flowchart of an image search method provided by an embodiment of the present application.

[0011] Figure 3 It is provided by an embodiment of the present application, and is a list of top-ranked images with the query keyword "lady" obtained based on the Semantic Alignment Network (SAN).

[0012] Figure 4 It is a flowchart of another image search method provided by an embodiment of the present application.

[0013] Figure 5 It is a flowchart of an image indexing method provided by an embodiment of the present application.

[0014] Figure 6 It is a flowchart of another image indexing method provided by an embodiment of the present application.

[0015] Figure 7 It is a schematic structural diagram of an image search device provided by an embodiment of the present application.

[0016] Figure 8 It is a schematic structural diagram of an image indexing device provided by an embodiment of the present application.

[0017] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0018] Most state-of-the-art text-to-photo retrieval works rely on encoding images into an embedding vector that can align the visual space and the word space. However, for retrieval tasks, relying on current word embeddings may be problematic because the existing word space is based on word co-occurrence information in the corpus rather than semantic similarity between words. For example, in current word embeddings, the cosine similarity between the two words "lady" and "gentleman" and "adult" is higher than the cosine similarity between either "lady" or "gentleman" and "adult". Therefore, performing text-to-photo search using the keyword "lady" based on the image vectors learned from current word embeddings will result in unexpected results, that is, the top-ranked images related to the keyword "lady" may be images related to "gentleman".

[0019] To solve the above problems, the present application provides an image search method, an image indexing method, and an electronic device, which improve the deficiencies of current word vectors by using a Semantic Alignment Network (SAN) and enhance the search accuracy of text-to-photo.

[0020] To facilitate the understanding of this application, the SAN provided by the embodiments of this application will be described in detail below.

[0021] Embodiments disclosed in this application will be described in detail below, and examples thereof are shown in the accompanying drawings. Among them, the same or similar reference numerals are used throughout to denote the same or similar elements or elements serving the same or similar functions. The embodiments described with reference to the accompanying drawings are merely exemplary and are intended to illustrate rather than limit this application.

[0022] The SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space. This visual-semantic space defines the mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0023] Semantic constraints refer to language constraints, for example, the language relationship between words. The semantic embedding mapped to a certain image embedding is generated based on semantic constraints. In other words, language constraints are injected into the word vectors mapped to the image embedding. Therefore, this can improve the usefulness of the text-to-photo search task.

[0024] In some embodiments, the semantic constraints may include synonym relationships and antonym relationships. A synonym relationship means that one word or term is synonymous with another word or term. For example, "man" and "adult", "woman" and "adult". An antonym relationship means that one word is the opposite of another word. For example, "man" and "woman".

[0025] Figure 1 is a framework diagram of a semantic alignment network (SAN) provided by the embodiments of this application. The SAN 100 includes a visual model, a language model, and an affix sub-network (WR-Net). The visual model can be used to extract the features of at least one image and convert the features of at least one image into at least one image embedding. The language model can be used to predict the labels of each image in at least one image to obtain a group of word vectors. The affix sub-network (WR-Net) can be used to convert the group of word vectors into semantic embeddings, thereby implementing a semantic embedding extractor. Therefore, the mapping relationship between the image embedding and the semantic embedding is generated.

[0026] The visual model can be used to extract features of at least one image and transform the features of the at least one image into at least one image embedding. Specifically, the visual model may include a Convolutional Neural Network (CNN). The Convolutional Neural Network (e.g., ResNet) is used as a feature extractor for images. The Convolutional Neural Network does not include a softmax prediction layer, but includes multiple convolutional filters (with skip connections), batch normalization, and pooling layers, and subsequent multiple fully connected neural network layers. The Convolutional Neural Network is trained through a softmax output layer to predict one out of 1000 object categories from a dataset (e.g., ILSVRC 2012 1K dataset). The output of the last global pooling layer of the Convolutional Neural Network is a 2048-dimensional vector and is used as the image embedding of the image. That is to say, this image embedding is the deep CNN feature of the image.

[0027] The visual model may further include a core part. The core part of the visual model is trained to predict the semantic embedding of each image through a projection layer and a similarity metric. The projection layer is a linear transformation that maps the 2048-dimensional deep vector to the same representation native to the language model.

[0028] The language model can be used to predict the label of each image in at least one image to obtain a group of word vectors. The language model can be the skip-gram text modeling architecture introduced by Mikolov. The label of each image can be unannotated text or vocabulary. The unannotated text or vocabulary may include multiple words or terms.

[0029] The skip-gram text modeling architecture introduced by Mikolov has been proven to be able to effectively learn semantically meaningful floating-point representations of terms from unannotated text. The skip-gram text modeling architecture learns to represent each word or term as an embedding vector with a fixed length by predicting adjacent terms in the unannotated text. These embedding vectors represent word vectors, also known as text embeddings.

[0030] In an example of the skip-gram text modeling architecture introduced by Mikolov, a 300-dimensional word vector (i.e., text embedding) is created to represent each word or term.

[0031] The Word Affix Sub-Network (WR-Net) can be used to transform the group of word vectors into semantic embeddings, thus implementing a semantic embedding extractor. Specifically, WR-Net uses synonym and antonym relationships extracted from general lexical resources or specific application ontologies to fine-tune the distributed word vectors.

[0032] The WR-Net may include two fully connected layers, a batch normalization layer, and ReLU. WR-Net can be defined as:

[0033] WR(w) = 2σ RELU ((M1w))

[0034] Wherein, M1 and M2 are two fully connected layers, BN(.) is a batch normalization layer, and σ RELU is the activation function of ReLU.

[0035] In some embodiments, the training of the SAN includes training the WR-Net according to a semantic-alignment loss generated by semantic constraints, so that the WR-Net is configured as the semantic embedding extractor; and includes training the visual model according to a visual-semantic loss, so as to provide the visual-semantic space.

[0036] Since the WR-Net is trained according to the semantic-alignment loss generated by semantic constraints, and the visual model is trained according to the visual-semantic loss, the entire SAN is trained. That is, in order to minimize the semantic-alignment loss and the visual-semantic loss, the stochastic gradient descent method (SGD) can be used to iteratively find the network parameters and train the entire network. Specifically, the WR-Net is trained alone until it converges, and then it is frozen and used as the semantic embedding extractor. Then, the visual-semantic loss is used as a guide for training the visual model, thereby training the entire network.

[0037] Furthermore, in some examples, the training of the WR-Net includes adjusting the word vector group according to the semantic-alignment loss to generate another word vector group as the semantic embedding.

[0038] For the word vector group W = {w1, w2, …, w n}, where each word in the text or vocabulary has a vector (i.e., a label). Semantic constraints, such as synonym relationships and antonym relationships, are injected into this vector space (i.e., the word vector group W), resulting in another word vector group W′ = {w′1, w′2, …, w′ n} based on the semantic-alignment loss. The word vector group W′ can also be referred to as the new semantic-alignment word vector or the new word vector group, while the word vector group W can be referred to as the original word vector group.

[0039] As described above, the semantic constraints can include synonym relationships and antonym relationships. Specifically, in some embodiments, the semantic constraints include: a first sub-constraint of the word vectors of a pair of synonyms, the word vectors of the pair of synonyms being adjacent to each other in the other word vector group; a second sub-constraint of the word vectors of a pair of antonyms, the word vectors of the pair of antonyms being separated from each other in the other word vector group; and a third sub-constraint of the other word vector group, the other word vector group retaining the information contained in the word vector group.

[0040] In the first sub-constraint, the word vectors of a pair of synonyms are pulled closer in another word vector group W'. For example, the word vectors of the pair of synonyms are adjacent to each other in another word vector group W'.

[0041] In the second sub-constraint, the word vectors of a pair of synonyms are pushed away from each other in another word vector group W'. For example, the word vectors of a pair of antonyms are separated from each other in another word vector group W'.

[0042] In the third sub-constraint, another word vector group retains the information contained in the word vector group. Since the synonym relationship and the antonym relationship are injected into the new representation (i.e., the word vector group W), the inferred word vectors need to be as close as possible to the original word vectors. In this case, another word vector group W' needs to retain the information contained in the word vector group W, that is, the inferred word vectors need to retain the information in the original word vectors.

[0043] Accordingly, the semantic-alignment loss includes a synonym loss, an antonym loss, and a space loss. Among them, the synonym loss is used to implement the first sub-constraint, the antonym loss is used to implement the second sub-constraint, and the space loss is used to implement the third sub-constraint.

[0044] In some examples, the semantic-alignment loss can be obtained when performing a predetermined operation on the synonym loss, the antonym loss, and the space loss. For example, performing an addition operation on the synonym loss, the antonym loss, and the space loss. Therefore, the semantic-alignment loss can be the sum of the synonym loss, the antonym loss, and the space loss. In another example, performing a weighted sum operation on the synonym loss, the antonym loss, and the space loss. Therefore, the semantic-alignment loss can be the weighted sum of the synonym loss, the antonym loss, and the space loss, as follows:

[0045] SAL = αSynL(W') + βAntL(W') + γSpaceL(W, W')

[0046] where SAL is the semantic-alignment loss, SynL(W') is the synonym loss, AntL(W') is the antonym loss, SpaceL(W, W') is the space loss, and α, β, and γ control the relative strength of these losses.

[0047] In some examples, the synonym loss is represented by the distance between the word vectors of a pair of synonyms in the synonym set. Further, the synonym loss SynL(W') is defined as:

[0048] SynL(W') = -∑ (a,b)∈s d(w' a , w' b )

[0049] Among them, d is a distance function, and cosine similarity is used to evaluate synonym pairs; a and b are a pair of synonyms, S is a synonym set with synonym pairs, w′ a and w′ b are the word vectors of word a and word b in another word vector group W′.

[0050] In some examples, the antonym loss is represented by the difference between the distance between the word vectors of a pair of antonyms in the antonym set and the minimum distance between the antonyms in the antonym set. Further, the antonym loss AntL(W′) is defined as:

[0051] AntL(W′) = Σ (a,b)∈A max(d(w′ a , w′ b ) - m, 0)

[0052] where d is a distance function, and cosine similarity is used to evaluate antonym pairs; a and b are a pair of antonyms, w′ a and w′ b are the word vectors of word a and word b in another word vector group W′, A is an antonym set with antonym pairs, and m is the margin or minimum distance between the antonyms in the antonym set.

[0053] In some examples, the space loss is represented by the distance between the word vector of a word in a word vector group and the word vector of the said word in another word vector group, and the distance between the word vector of the said word in another word vector group and the word vector of the neighboring word of the said word in another word vector group.

[0054] Further, the space loss SpaceL(W, W′) is defined as:

[0055] SpaceL(W, W′) = -∑ i [d(w′ i , w i ) + ∑ j∈N(i) d(w′ i , w′ j )]

[0056] where d is a distance function, d(w′ i , w i ) is the distance between the word vector of word i in the word vector group W (i.e., the original space or W space) and the word vector of word i in another word vector group W′ (i.e., the new space or W′ space), N(i) is the neighboring word of word i in the original space (i.e., W space), and d(w′ i , w′ j) is the distance between the word vector of word i in another word vector group W′ and the word vector of the neighboring word j of word i in another word vector group W′.

[0057] As described above, the visual model is trained according to the visual-semantic loss to provide a visual-semantic space. In some embodiments of the present application, a combination of dot product similarity and triplet loss is used for the semantic-alignment loss. That is, in embodiments where the semantic-alignment loss can be a weighted sum of synonym loss, antonym loss, and spatial loss, the visual model is trained such that the dot product similarity between the output of the visual model and the semantic embedding of the correct label is higher than the dot product similarity between the output of the visual model and other randomly selected words or terms.

[0058] Specifically, in some embodiments, the visual-semantic loss is represented by the distance between the depth vector and each word vector in the word vector group and the distance between the depth vector and the word vector of the correct (Ground-truth) label. Further, the visual-semantic loss (VS Loss) of each training instance is defined as follows:

[0059] VS Loss = Σ j≠labeL max(d(I, wr(W j )) - d(I, WR(w label )) + m vs , 0)

[0060] where I is the depth vector, d is the distance function, d(I, WR(w j )) is the distance between the depth vector I and the word vector w in the word vector group W j , w label is the word vector of the correct (Ground-truth) label, d(I, WR(w label )) is the distance between the depth vector I and the word vector w of the correct label label , m vs is the margin distance.

[0061] Figure 2 is a flowchart of an image search method provided by an embodiment of the present application. This method can be executed by an electronic device, including but not limited to a computer, a server, etc. This method includes the following actions / operations.

[0062] 210: Obtain a query keyword.

[0063] The query keyword can be information included in the image, for example, "man", "woman", "adult", etc., or the location where the image is captured.

[0064] 220: Obtain the target semantic embedding of the query keyword through the Semantic Alignment Network (SAN), and search for at least one target image corresponding to the query keyword according to the target semantic embedding by the SAN.

[0065] The SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space that defines the mapping relationship between at least one image embedding and semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0066] Through the SAN, input the query keyword into the SAN to obtain the target semantic embedding of the query keyword, and find at least one target image corresponding to the query keyword according to the target semantic embedding in the visual-semantic space of the SAN. For example, when the query keyword is "lady", as Figure 3 shown, the top-ranked images are all "lady" images, rather than "man" images, which indicates that the connection between the visual space and the word space has a better vector representation and the result is more accurate.

[0067] In these embodiments, the SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space that defines the mapping relationship between at least one image embedding and semantic embedding, and each semantic embedding is generated based on semantic constraints. Obtain the target semantic embedding of the query keyword through the SAN, and find at least one target image corresponding to the query keyword according to the target semantic embedding in the visual-semantic space of the SAN, thereby improving the deficiencies of the current word vectors and enhancing the search accuracy from text to photo.

[0068] In some embodiments, for obtaining the target semantic embedding of the query keyword in step 220, first predict the query keyword through a language model to obtain the word vector of the query keyword, and then convert the word vector of the query keyword into a target semantic embedding through the WR-Net.

[0069] As described above, find at least one target image corresponding to the query keyword based on the target semantic embedding. In some examples, the at least one target image may be the nearest image obtained in the visual-semantic space of the SAN based on the target semantic embedding, for example, the images arranged in ascending order of their respective distances from the target semantic embedding. That is, the at least one target image is the top-ranked image in the visual-semantic space of the SAN. In addition, the at least one target image may include the image with the shortest distance from the target semantic embedding in the visual-semantic space.

[0070] Figure 4It is a flowchart of another image search method provided by an embodiment of the present application. This method can be executed by an electronic device, including but not limited to a computer, a server, etc. Based on the above embodiment, this method further includes the following actions / operations.

[0071] 410: Obtain at least one image.

[0072] The at least one image can be obtained from the user's album in the electronic device or can be taken on-site.

[0073] 420: Convert the at least one image into the at least one image embedding through the SAN, thereby defining a mapping relationship in the visual-semantic space.

[0074] Through the SAN, input at least one image into the SAN, and then convert it into at least one image embedding, thereby defining a mapping relationship in the visual-semantic space.

[0075] In these embodiments, through the SAN configured as a semantic embedding extractor, at least one image is input into the SAN and then converted into at least one image embedding, thereby defining a mapping relationship in the visual-semantic space. Therefore, the image is mapped to a finer semantic space, improving the accurate indexing of the image and helping to enhance the search accuracy from text to photo.

[0076] In some embodiments, for converting at least one image into at least one image embedding through the SAN in step 420, first obtain the depth features of the at least one image through the visual model of the SAN, and then convert the depth features of the at least one image into at least one image embedding.

[0077] Figure 5 It is a flowchart of an image indexing method provided by an embodiment of the present application. This method can be executed by an electronic device, including but not limited to a computer, a server, etc. This method includes the following actions / operations.

[0078] 510: Obtain at least one image.

[0079] The at least one image can be obtained from the user's album in the electronic device or can be taken on-site.

[0080] 520: Convert the at least one image into at least one image embedding through the SAN, thereby providing a visual-semantic space.

[0081] This visual-semantic space defines the mapping relationship between at least one image embedding and semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0082] Through the SAN, at least one image is input into the SAN and then converted into at least one image embedding, thereby providing a visual-semantic space.

[0083] In these embodiments, through the SAN configured as a semantic embedding extractor, at least one image is input into the SAN and then converted into at least one image embedding, thereby defining a mapping relationship in the visual-semantic space. Thus, the image is mapped to a finer semantic space, improving the accurate indexing of the image and contributing to enhancing the search accuracy from text to photo.

[0084] In some embodiments, for the conversion of at least one image into at least one image embedding through the SAN in step 520, first, the depth features of the at least one image are obtained through the visual model of the SAN, and then the depth features of the at least one image are converted into at least one image embedding.

[0085] Figure 6 It is a flowchart of another image indexing method provided by an embodiment of the present application. This method can be executed by an electronic device, including but not limited to a computer, a server, etc. Based on the above embodiments, this method further includes the following actions / operations.

[0086] 610: Obtain a query keyword.

[0087] The query keyword can be information contained in the image, for example, "man", "woman", "adult", etc., or the location where the image is captured.

[0088] 620: Obtain the target semantic embedding of the query keyword through the SAN, and search for at least one target image corresponding to the query keyword through the SAN according to the target semantic embedding.

[0089] The SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space. This visual-semantic space defines the mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0090] Through the above SAN, the query keyword is input into the SAN to obtain the target semantic embedding of the query keyword, and at least one target image corresponding to the query keyword is found according to the target semantic embedding in the visual-semantic space of the SAN. For example, when the query keyword is "woman", in the above SAN, as Figure 3 shown, the top-ranked images are all "woman" images, rather than "man" images, which indicates that there is a better vector representation of the connection between the visual space and the word space, and the result is more accurate.

[0091] In these embodiments, the SAN is configured as a semantic embedding extractor and is used to provide a vision-semantic space that defines the mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints. The target semantic embedding of the query keyword is obtained through the SAN, and at least one target image corresponding to the query keyword is found according to the target semantic embedding in the vision-semantic space of the SAN, thus improving the deficiencies of the current word vectors and enhancing the search accuracy from text to photo.

[0092] In some embodiments, in order to obtain the target semantic embedding of the query keyword in step 620, first, the query keyword is predicted by a language model to obtain the word vector of the query keyword, and then the word vector of the query keyword is converted into a target semantic embedding through the WR-Net.

[0093] As described above, at least one target image corresponding to the query keyword is found based on the target semantic embedding. In some examples, the at least one target image may be the nearest image obtained in the vision-semantic space of the SAN based on the target semantic embedding. For example, the images arranged in ascending order of their respective distances from the target semantic embedding. That is, the at least one target image is the image ranked at the top in the vision-semantic space of the SAN. In addition, the at least one target image may include the image with the shortest distance from the target semantic embedding in the vision-semantic space.

[0094] Figure 7 It is a schematic structural diagram of an image search device provided by an embodiment of the present application. The device 700 may include a first acquisition module 710 and a second acquisition module 720.

[0095] The first acquisition module 710 is available for obtaining a query keyword. The second acquisition module 720 is available for obtaining the target semantic embedding of the query keyword through a semantic alignment network (SAN), and searching for at least one target image corresponding to the query keyword according to the target semantic embedding through the SAN. Wherein, the SAN is configured as a semantic embedding extractor and is used to provide a vision-semantic space, which defines the mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0096] It should be noted that the above description of the image search method in the above embodiments also applies to the device of the exemplary embodiments of the present application and will not be described herein.

[0097] Figure 8 It is a schematic structural diagram of an image indexing device provided by an embodiment of the present application. The device 800 may include an acquisition module 810 and a conversion module 820.

[0098] The acquisition module 810 can be used to obtain at least one image. The conversion module 820 can be used to convert at least one image into at least one image embedding through a SAN, thereby providing a visual-semantic space. The visual-semantic space defines a mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints.

[0099] It should be noted that the above description of the image search method in the above embodiments also applies to the devices in the exemplary embodiments of the present application, and will not be described herein.

[0100] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 900 may include a processor 910 and a memory 920 coupled together.

[0101] The memory 920 is configured to store executable program instructions. The processor 910 may be configured to read the executable program instructions stored in the memory 920 to execute a program corresponding to the executable program instructions, thereby executing the image search method described in the above embodiments or any non-conflicting combination of the above embodiment methods, or executing the image indexing method described in the above embodiments or any non-conflicting combination of the above embodiment methods.

[0102] In one example, the electronic device 900 may be a computer, a server, etc. In another example, the electronic device 900 may be an independent component integrated in a computer or a separator.

[0103] The present application further provides a non-transitory computer-readable storage medium, which may be in the memory 920. The non-transitory computer-readable storage medium stores instructions. When the instructions are executed by the processor, the processor executes the method as described in the foregoing embodiments.

[0104] Those skilled in the art can understand that, in combination with the examples described in the embodiments disclosed in this specification, the units and algorithm steps can be implemented by electronic hardware, computer software, or a combination thereof. To clearly describe the interchangeability between hardware and software, the composition and steps of each embodiment are generally described according to functions above. Whether these functions are implemented by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but should not consider this implementation manner to exceed the scope of the present application.

[0105] Those skilled in the art can clearly understand that, for the convenience and brief description, for the detailed working processes of the foregoing systems, devices, and units, reference may be made to the corresponding processes in the method embodiments, and the details will not be repeated herein.

[0106] In the multiple embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the described device embodiments are merely exemplary. For example, the unit division is only a logical function division and may be other divisions in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some functions can be ignored or not executed. In addition, the mutually coupled, directly coupled, or communication-connected relationships shown or discussed can be implemented through some interfaces. The indirect coupling or communication connection between devices or units can be achieved in electronic, mechanical, or other forms.

[0107] The units described as independent components may or may not be physically separated. The components shown as units may or may not be physical units, may be located in one place, or may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to implement the solutions of the embodiments of this application.

[0108] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0109] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, basically, or the part that contributes to the prior art, or all or part of the technical solution, can be implemented in the form of a software product. The computer software product is stored in a storage medium, for example, a non-transitory computer-readable storage medium, and includes several instructions for instructing a computer device (which can be a personal computer, a server, or a network device) to execute all or part of the steps of the method described in the embodiments of this application. The storage medium includes any medium that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0110] The above description is only specific embodiments of this application, but is not intended to limit the protection scope of this application. Any equivalent modifications or substitutions conceived by those skilled in the art within the technical scope of this application should fall within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An image search method, characterized in that, including: obtaining a query keyword; and obtaining a target semantic embedding of the query keyword through a Semantic Alignment Network (SAN), and searching for at least one target image corresponding to the query keyword according to the target semantic embedding through the SAN; wherein, the SAN is configured as a semantic embedding extractor and is used to provide a visual-semantic space; the visual-semantic space defines a mapping relationship between at least one image embedding and a semantic embedding, and each semantic embedding is generated based on semantic constraints; the SAN includes: a visual model for extracting features of at least one image and converting the features of the at least one image into the at least one image embedding; a language model for predicting labels of each image in the at least one image to obtain a word vector group; and an affix sub-network (WR-Net) for converting the word vector group into the semantic embedding, thereby implementing the semantic embedding extractor.

2. The method according to claim 1, wherein the training of the SAN includes: training the WR-Net according to a semantic-alignment loss generated by the semantic constraints, so that the WR-Net is configured as the semantic embedding extractor; and training the visual model according to a visual-semantic loss, thereby providing the visual-semantic space.

3. The method according to claim 2, characterized in that, the training of the WR-Net includes: adjusting the word vector group according to the semantic-alignment loss to generate another word vector group as the semantic embedding.

4. The method according to claim 3, characterized in that the semantic constraints include: a first sub-constraint of word vectors of a pair of synonyms, and the word vectors of the pair of synonyms are adjacent to each other in the another word vector group; a second sub-constraint of word vectors of a pair of antonyms, and the word vectors of the pair of antonyms are separated from each other in the another word vector group; and a third sub-constraint of the another word vector group, and the another word vector group retains information included in the word vector group; wherein, the semantic-alignment loss includes a synonym loss, an antonym loss and a space loss; the synonym loss is used to implement the first sub-constraint, the antonym loss is used to implement the second sub-constraint, and the space loss is used to implement the third sub-constraint.

5. The method according to claim 4, characterized in that the synonym loss is represented by the distance between the word vectors of the pair of synonyms in the synonym set.

6. The method according to claim 4, wherein the antonym loss is represented by the difference between the distance between the word vectors of the pair of antonyms in the antonym set and the minimum distance between the antonyms in the antonym set.

7. The method according to claim 4, wherein the space loss is represented by a first distance and a second distance; the first distance is the distance between the word vector of a word in the word vector group and the word vector of the word in the another word vector group; the second distance is the distance between the word vector of the word in the another word vector group and the word vector of an adjacent word of the word in the another word vector group.

8. The method according to claim 2, wherein the visual-semantic loss is represented by the distance between the depth vector and each word vector in the word vector group and the distance between the depth vector and the word vector of the correct (Ground-truth) label.

9. The method according to claim 1, characterized in that, The obtaining of the target semantic embedding of the query keyword includes: Predicting the query keyword by the language model to obtain a word vector of the query keyword; and Converting the word vector of the query keyword into the target semantic embedding by the WR-Net.

10. The method according to claim 9, characterized in that, The at least one target image includes an image having the shortest distance from the target semantic embedding in the visual-semantic space.

11. The method according to claim 1, wherein It further includes: Obtaining at least one image; And Converting the at least one image into the at least one image embedding by the SAN, thereby defining the mapping relationship in the visual-semantic space.

12. The method according to claim 11, wherein The converting of the at least one image into the at least one image embedding by the SAN includes: Obtaining the depth features of the at least one image by the visual model of the SAN; and Converting the depth features of the at least one image into the at least one image embedding.

13. An image indexing method, characterized in that, It includes: Obtaining at least one image; And Converting the at least one image into at least one image embedding by a Semantic Alignment Network (SAN), thereby providing a visual-semantic space; wherein, the visual-semantic space defines the mapping relationship between the at least one image embedding and the semantic embedding, and each semantic embedding is generated based on semantic constraints; The SAN includes: A visual model for extracting the features of the at least one image and converting the features of the at least one image into the at least one image embedding; A language model for predicting the labels of each image in the at least one image to obtain a group of word vectors; and An Affix Sub-Network (WR-Net) for converting the group of word vectors into the semantic embedding, thereby implementing a semantic embedding extractor.

14. The method according to claim 13, wherein The training of the SAN includes: Training the WR-Net according to a semantic-alignment loss generated by the semantic constraints, so that the WR-Net is configured as the semantic embedding extractor; and Training the visual model according to a visual-semantic loss, thereby providing the visual-semantic space.

15. The method according to claim 14, wherein The training of the WR-Net includes: Adjusting the group of word vectors according to the semantic-alignment loss to generate another group of word vectors as the semantic embedding.

16. The method according to claim 15, wherein The semantic constraints include: A first sub-constraint of word vectors of a pair of synonyms, and the word vectors of the pair of synonyms are adjacent to each other in the another group of word vectors; A second sub-constraint of word vectors of a pair of antonyms, and the word vectors of the pair of antonyms are separated from each other in the another group of word vectors; and A third sub-constraint of the another group of word vectors, and the another group of word vectors retains the information included in the group of word vectors; Wherein, the semantic-alignment loss includes a synonym loss, an antonym loss and a space loss; the synonym loss is used to implement the first sub-constraint, the antonym loss is used to implement the second sub-constraint, and the space loss is used to implement the third sub-constraint.

17. The method according to claim 16, wherein The synonym loss is represented by the distance between the word vectors of the pair of synonyms in the synonym set.

18. The method according to claim 16, characterized in that, The antonym loss is represented by the difference between the distance between the word vectors of the pair of antonyms in the antonym set and the minimum distance between the antonyms in the antonym set.

19. The method according to claim 16, wherein The spatial loss is represented by a first distance and a second distance; the first distance is the distance between the word vector of a word in the word vector group and the word vector of the word in the other word vector group; the second distance is the distance between the word vector of the word in the other word vector group and the word vector of the neighboring word of the word in the other word vector group.

20. The method according to claim 14, wherein The vision-semantic loss is represented by the distance between the depth vector and each word vector in the word vector group and the distance between the depth vector and the word vector of the correct (Ground-truth) label.

21. The method according to claim 13, wherein The converting the at least one image into at least one image embedding includes: obtaining depth features of the at least one image through the vision model; and converting the depth features of the at least one image into the at least one image embedding.

22. The method according to claim 13, characterized in that, Further includes: obtaining a query keyword; obtaining a target semantic embedding of the query keyword through the SAN, and searching for at least one target image corresponding to the query keyword according to the target semantic embedding through the SAN.

23. The method according to claim 22, wherein The obtaining the target semantic embedding of the query keyword includes: predicting the query keyword through the language model to obtain the word vector of the query keyword; and converting the word vector of the query keyword into the target semantic embedding through the WR-Net.

24. The method according to claim 23, characterized in that The at least one target image includes an image having the shortest distance from the target semantic embedding in the vision-semantic space.

25. An electronic device, characterized in that, Includes a processor and a memory; the memory stores instructions, and when the instructions are executed by the processor, the processor executes the method described in any one of claims 1-12 and claims 13-24.

26. A non-transitory computer-readable storage medium, characterized in that, Stores instructions; when the instructions are executed by a processor, the processor executes the method described in any one of claims 1-12 and claims 13-24.

Citation Information

Patent Citations

  • Image search system and method for personalized photo applications using semantic networks

    US20150026101A1