Text-to-Visual Machine Learning Embedding Technology
By generating training data sets and optimizing models using loss function, the accuracy and efficiency problems of digital image search systems in the prior art are solved when processing complex text queries, and more efficient image search results are achieved.
Patent Information
- Application Number
- CN202010182685.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-30
- Filing Date
- 2020-03-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-03-16
AI Technical Summary
Existing digital image search systems are prone to errors in text-based searches, especially when text queries contain a large amount of text, lack flexibility and accuracy, resulting in users needing to browse a large number of search results to find images of interest, which is inefficient.
By generating training data sets, including positive and negative digital image samples and associated text, the model is trained using machine learning, and the loss function is used to process the distance between positive and negative image embedding and text embedding separately, improving the accuracy and computational efficiency of the model.
It improves the accuracy and efficiency of digital image search, reduces the time for users to browse search results, and improves the model's discernment ability when processing complex text queries.
Smart Images

Figure CN112015940B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of digital images, and more particularly to digital image machine learning embedding techniques. Background Art
[0002] To return accurate search results, digital image search systems face many technical challenges, especially in instances involving text-based search. For example, conventional digital image search systems rely on image tags associated with digital images, which can be manually specified or automatically inferred, e.g., using machine learning-based image tagging techniques. Thus, to perform a search, the text included in a text query is matched to the tags associated with the digital image. However, these conventional systems and techniques are error-prone, especially when the text query includes a large amount of text, and typically due to the lack of the ability to support flexible variations in language descriptions.
[0003] In a conventional example, a text query including the text "a person sitting on a chair holding a dog by the beach" is received. Conventional digital image search systems based on tags typically return search results having any one of the tags also included in the text query. Thus, users of these conventional systems typically face search results that only include a person, a dog (a dog can include a canine or food), a chair, or a beach. The search results are also mixed with sub-combinations of these tags, e.g., a person with a dog, a person eating a hot dog, a chair on the beach, etc. Further, some conventional systems (e.g., inventory image-based search systems) may not even be able to return results due to the length of the text query. Thus, conventional digital image search systems may require users to manually navigate through hundreds of digital images in the search results to find the digital image of interest, may force users to initiate a large number of searches with different text combinations, etc. This results in user frustration due to the inefficiency of browsing and the inefficiency of the digital image search system's use of network and computing resources for transmitting, performing, and repeating these searches. Summary of the Invention
[0004] Describes text-to-visual machine learning embedding techniques that overcome challenges in conventional techniques in various ways. In one example, this is done by using a training data generation module to generate training data that improves the accuracy of a model trained using machine learning. For example, query-based training data can be generated based on text queries used to initiate a search for digital images and select digital images from search results. In this way, the association between text queries and digital images can be determined for a large number of digital images and text. The use of query-based training data can also be extended by using caption-based training data as part of multi-task learning, which improves training accuracy by limiting noise in the query-based training data and supports the use of long text sequences.
[0005] The training data generation module is also configured to generate negative digital image samples that improve accuracy when using machine learning to train a model. This is done by generating negative digital image samples that have a similar semantic and / or visual meaning to positive digital image samples but do not have exactly the same components as the positive digital image samples.
[0006] In one example, this is done by selecting negative digital image samples from a subset of digital images that have more than one text item, the text items do not include stop words, and are also included in the text associated with the positive digital image samples. In another example, this is done by selecting negative digital image samples from a subset of digital images that do not have every text item, the text items do not include stop words, and are also included in the text associated with the positive digital image samples. Then, this training data can be used to generate a model that supports a single unified text and digital image embedding space, which is configured to treat text and digital images as the same entity, and thus, overcomes the limitations of conventional techniques that are based only on text.
[0007] Also described is a machine learning training module that uses a loss function to train a model. Compared to conventional loss functions, this loss function supports improved accuracy and computational efficiency by separately processing the loss calculated between a positive image embedding generated from a positive digital image sample and a text embedding based on the text associated with the positive digital image sample from the loss calculated between a negative image embedding generated from a negative digital image sample and the text embedding. This allows the distance between the positive image embedding and the text embedding to decrease over time (during training), while the distance between the negative image embedding and the text embedding increases, thus improving model accuracy compared to conventional training techniques.
[0008] The following presents a simplified selection of concepts that are further described below in the Detailed Description. Similarly, the Summary of the Invention is not intended to identify key features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The detailed description is described with reference to the accompanying drawings. Entities represented in the figures may represent one or more entities, and thus, the singular or plural forms of the entities may be referred to interchangeably in the discussion.
[0010] Figure 1 is an illustration of a digital media environment operable to employ text-to-visual machine learning embedding techniques described herein in an example implementation.
[0011] Figure 2 depicts a system in an example implementation where a service provider system generates a query-based training dataset based on a text query and a digital image associated with the text query.
[0012] Figure 3 is a flowchart depicting a process in an example implementation where a training dataset is used to train a model using machine learning, the training dataset is generated based on digital images, and the text query is used to locate the digital images as part of a search.
[0013] Figure 4 depicts a system in an example implementation where a training dataset is generated that includes negative digital image samples selected based on positive digital image samples and associated text.
[0014] Figure 5 is a flowchart depicting a process in an example implementation where negative digital image samples are generated based on a comparison of text associated with the negative digital image samples and text associated with the positive digital image samples.
[0015] Figure 6 depicts a system in an example implementation that shows a machine learning training module using multi-task training to multi-task train a model based on the following training datasets: a query-based training dataset and a caption-based training dataset.
[0016] Figure 7 depicts a system that more particularly shows the operations of a machine learning training module in generating embeddings and using a loss function.
[0017] Figure 8 depicts in more detail Figure 7 a system showing the following operations of the text embedding module shown: generating a text embedding from text associated with a positive digital image sample.
[0018] Figure 9 Depicts a graphical comparison between a conventional triplet loss and a positive-aware triplet ranking loss as described herein.
[0019] Figure 10 Depicts a process in an example implementation where a model is trained based on a loss function that separately addresses the loss between a text embedding and a positive image embedding from the loss between the text embedding and a negative image embedding.
[0020] Figure 11 Illustrates an example system that includes various components of an example device that can be implemented as any type of computing device described and / or utilized herein to implement embodiments of the techniques described herein. Figures 1 to 10 described and / or utilized to implement embodiments of the techniques described herein. Detailed Description
[0021] Overview
[0022] To return accurate search results, digital image search systems face many technical and computational challenges, particularly in instances involving text-based search. To perform a search using a conventional digital image search system, the text included in a text query is matched to the labels associated with digital images. However, these conventional systems and techniques are error-prone, particularly when the text query includes a large amount of text. This is typically due to the lack of ability of conventional systems to support the flexibility of variations in language descriptions (e.g., "hot dog" as a food item and "hot dog" as a panting dog) and the lack of ability to address how to order text.
[0023] Accordingly, conventional image search systems may require a user to browse through hundreds of digital images in search results to find a digital image of interest, may force the user to initiate a large number of searches with different text combinations, and so on. This leads to user frustration due to the inefficiency of browsing and the inefficiency of the use of network and computational resources for transmitting and performing these searches. This challenge is further exacerbated by the reliability of the text used to express the text query in matching the underlying meaning of the text of the labels used to identify the images, which can be difficult to achieve in some instances (e.g., when describing the emotions caused by a scene in a digital image).
[0024] Accordingly, text-to-visual (i.e., semantic / visual) machine learning embedding techniques are described that overcome the challenges in conventional systems and techniques. This includes techniques for generating training data and training techniques for loss functions that can be used to support mapping digital images and text into a single unified embedding space and support overcoming conventional challenges.
[0025] The training data generation module uses multiple digital images and associated text to generate a training data set. In this instance, the associated text includes text queries that are used to locate the corresponding digital images (e.g., as part of an image search in a search engine, an inventory image system, etc.). In this way, the training data generation module can collect a large number of digital images that users select as corresponding to the text used to locate those digital images in an efficient manner. This overcomes the challenges in conventional techniques where the availability of accurate training data is limited (e.g., a limited number of samples) and expensive because it typically involves manual labeling, and manual labeling can lead to inaccuracies due to inconsistent application of the labels.
[0026] The training data generation module can also be configured to generate a title-based training data set, e.g., to support multi-task training as well as a query-based training data set. The title-based training data set includes digital images and titles associated with the digital images (e.g., image caption bars). The title-based training data set is used to address longer sentences and remove user query noise from the query-based training data set (e.g., for "clicked images" that do not correspond to text queries). The multi-task training achieved by using the title-based training data set with the query-based training data set improves the accuracy and computational efficiency of model training as part of machine learning as further described in the discussion below.
[0027] The training data generation module can also employ techniques for generating negative digital image samples. In machine learning as implemented by a machine learning training module, positive digital image samples are used as instances of "correct" correspondence with text, while negative digital image samples are used to improve, for example, the discrimination of a model selected in the following way: the negative digital image samples do not belong to the same category as the positive digital image samples. This is done by generating negative digital image samples that have a similar semantic meaning and / or visual meaning to the positive digital image samples but do not have exactly the same components as the positive digital image samples.
[0028] In one example, this is done by the training data generation module by selecting negative digital image samples from a subset of digital images that have more than one text item, the text items do not include stop words, and are also included in the text associated with the positive digital image samples. In another example, this is done by selecting negative digital image samples from a subset of digital images that do not have each text item, the text items do not include stop words, and are also included in the text associated with the positive digital image samples.
[0029] As part of using machine learning to train a model, the machine learning training module can also implement a loss function that further improves the accuracy and computational efficiency of the model. Continuing with the above example, the machine learning training module uses positive digital image samples, negative digital image samples, and text associated with the positive digital image samples to train the model. The machine learning training module uses machine learning to generate text embeddings from the text, e.g., via a recurrent neural network. Positive image embeddings are also generated from the positive digital image samples and negative image embeddings are generated from the negative digital image samples, e.g., via a convolutional neural network encoder.
[0030] In this example, the loss function is configured to evaluate the loss between the text embedding and the positive image embedding separately from the loss between the text embedding and the negative image embedding. This results in the positive image embeddings having increased similarity (and thus, clustering together) relative to the text embedding, and the negative image embeddings having increased dissimilarity relative to the text embedding. This improves the model's ability to discriminate between these samples, i.e., improves the model accuracy. In this way, the model accuracy is improved compared to conventional loss functions that do not support the ability to separately address these losses.
[0031] In the following discussion, an example environment in which the techniques described herein can be employed is first described. Example processes that can be performed in the example environment as well as other environments are also described. Thus, the execution of the example processes is not limited to the example environment, and the example environment is not limited to the execution of the example processes.
[0032] Example Environment
[0033] Figure 1 FIG. 13 is a diagram of a digital media environment 100 operable in an example implementation to employ the text-to-visual machine learning embedding techniques described herein. The illustrated environment 100 includes a service provider system 102 and a client device 104 communicatively coupled via a network 106. The service provider system 102 and the client device 104 can be implemented using a variety of different configurations of computing devices.
[0034] For example, a computing device can be configured as a desktop computer, a laptop computer, a mobile device (e.g., assuming a handheld configuration such as a tablet computer or a mobile phone as illustrated for client device 104), etc. Thus, the range of computing devices can span from full-resource devices with substantial memory and processor resources (e.g., personal computers, gaming consoles) to low-resource devices with limited memory and / or processing resources (e.g., mobile devices). Additionally, a computing device can represent multiple different devices, such as multiple servers utilized by an enterprise to perform operations "in the cloud" as illustrated for service provider system 102 and as further described with respect to Figure 11 and further described below.
[0035] The client device 104 as illustrated includes a communication module 108 (e.g., a browser or a network-enabled application), which can be executed using a computer-readable storage medium and a processing system to access the functionality of the service provider system 102 via the network 106. This functionality can take various forms, such as using a search module 114 to search for digital images 110 illustrated as being stored in a storage device 112. Other examples of features that can be supported by the functionality described herein include machine translation, text retrieval, speech recognition, text summarization, etc. Further, although the functionality is illustrated as being implemented by the service provider system 102, the functionality can be implemented wholly or partially locally by the client device 104.
[0036] For example, the communication module 108 can receive a text query 116—“running shoes” via a user interface 118. The text query 116 is transmitted via the network 106 and processed by the search module 114. The search module 114 implements a single unified text and digital image embedding space 122 using a model 120 trained using machine learning to perform the search. This single unified text and digital image embedding space 122 overcomes the limitations of conventional text-only embedding techniques when they are used to address the relationship between text and digital images (e.g., to obtain a visual intuition of “what” is expressed in the text).
[0037] As previously described, conventional search techniques are error-prone, especially when the text query includes a large amount of text. This is typically due to the lack of the ability to support the flexibility of variations in language descriptions and variations of language descriptions for different objects. In the illustrated example, for instance, a conventional system may match the text query “running shoes” with digital images having labels of any pivot (i.e., text terms that are not stop words and are used as the basis for performing the search) with text, examples of which include a digital image 124(1) of a running dog, a digital image 124(3) of a shoe, and a digital image 124(4) of a person running along the actual search target, e.g., a digital image 124(3) of running shoes. Stop words are common words that are filtered as being irrelevant to the search, such as “and”, “the”, “a”, “an”, etc. as general words.
[0038] However, in the techniques described herein, a single unified text and digital image embedding space 122 is generated for the model 120 as part of machine learning, which overcomes these challenges with improved accuracy and computational efficiency compared to conventional techniques. For example, using the model 120 to search for “gold bowl” will limit and even eliminate several instances of digital images including goldfish, gold ingots, etc., which are commonly encountered in conventional techniques.
[0039] To this end, the digital media environment 100 described herein implements various functionalities that can be performed together or in sub - combinations as further described in the corresponding sections in the following discussion. In the illustrated example, the service provider system 102 employs a training data generation module 126 to generate a training data set 128 that is used by a machine learning training module 130 to train a model 120 using a loss function 132. The training data set 128 can be based on various different types of text that can be associated with digital images.
[0040] In the Query - based Training Data Set section following the subsequent discussion, the training data set 128 is generated by the training data generation module 126 using multiple digital images and associated text. In this instance, the associated text includes text queries that are used to locate the corresponding digital images. For example, the training data generation module 126 can receive data including a text query (e.g., running shoes) and a digital image, such as digital image 124(3), included in a user - selected digital image search result. In this way, the training data generation module 126 can efficiently collect a number of digital images and text that the user selects as corresponding to those digital images. This overcomes challenges in conventional techniques where the availability of accurate training data is limited (e.g., a limited number of samples) and expensive because it typically involves manual labeling, which can lead to inaccuracies due to inconsistent application of the labels.
[0041] The training data generation module 126 can also be configured to generate a caption - based training data set as part of the training data set 128 as also described in the Query - based Training Data Set section, e.g., as part of multi - task training. The caption - based training data set includes digital images and captions associated with the digital images (e.g., image caption bars). For example, the caption - based training data set can be used in combination with the query - based training data set to train the model 120 to address longer sentences, text sequences, and remove user query noise from the query - based training data set, e.g., for "clicked images" that do not correspond to text queries. Using the caption - based training data set along with the query - based training data set improves the accuracy and computational efficiency of the model 120 as further described in the corresponding sections.
[0042] As part of generating the training data set 128, the training data generation module 126 can also employ techniques for generating negative digital image samples. In machine learning, as implemented by the machine learning training module 130, positive digital image samples are used as instances of "correct" correspondence with text, while negative digital image samples are used to improve, for example, the discriminative power of the model 120 selected in the following manner: negative digital image samples do not belong to the same category as positive digital image samples.
[0043] As further described in the negative digital image sample generation section, the training data generation module 126 can automatically and without user intervention generate negative digital image samples in a manner that improves the accuracy of the model 120. To this end, the training data generation module 126 selects positive digital image samples from multiple digital images with associated text (e.g., the text queries or captions described above).
[0044] In one instance, the training data generation module 126 generates a subset from multiple digital images that includes digital images that do not have any item associated with the associated text of the positive digital image sample. For example, assume that the positive digital image sample has the associated text "person on a motorcycle". Then, the digital images are filtered to form a subset of digital images that are not associated with "person" or "motorcycle". Then, the subset is used to select negative digital image samples. For example, the training data generation module 126 can automatically and without user intervention select the digital image in the subset that is "closest" (by comparing the corresponding embeddings) to the positive digital image sample, e.g., the smallest squared distance. For example, this can be performed on the query-based training data described above. In this way, negative digital image samples can improve the ability of the model 120 to distinguish between "good" and "bad" examples of digital image and text associations.
[0045] In another example, the training data generation module 126 can automatically generate even "harder" negative digital image samples without user intervention. To this end, in this example, the training data generation module 126 also generates a subset from the following plurality of digital images: the plurality of digital images do not include the digital images having each item, and the items do not include stop words (i.e., "center points") in the text associated with the positive digital image samples. Then, the training data generation module 126 selects negative digital image samples from this subset. For example, again assuming that the positive digital image sample has the associated text of "person on a motorcycle". Then, the digital images having "person" or "motorcycle" are filtered out from the plurality of digital images, and the remaining digital images form a subset. Then, the subset is used to select negative digital image samples. For example, this can be performed on title-based training data that typically includes a large amount of text as described above. As a result, the model 120 is also able to distinguish between "good" and "bad" examples of digital image and text associations as part of the training.
[0046] As described in the loss function section in the following discussion, the machine learning training module 130 can also implement a loss function 132 as part of training and using the model 120, and this loss function 132 further improves the accuracy and computational efficiency of the model 120. Continuing with the above example, the machine learning training module 130 uses the positive digital image samples, negative digital image samples, and the text associated with the positive digital image samples to train the model 120. The machine learning training module 130 uses machine learning to generate text embeddings from the text, for example, through a recurrent neural network. Positive image embeddings are also generated from the positive digital image samples and negative image embeddings are generated from the negative digital image samples, for example, through a convolutional neural network encoder.
[0047] In this example, the loss function 132 is configured to evaluate the loss between the text embedding and the positive image embedding separately from the loss between the text embedding and the negative image embedding. This enables the positive image embeddings to have improved similarity (and thus, be clustered together) with respect to the text embedding during training, and the negative image embeddings to have improved dissimilarity with respect to the text embedding, for example, to make the clusters "tighter". This improves the ability of the model 120 to distinguish these samples, i.e., improves the accuracy of the model 120. In this way, the accuracy of the model is improved as compared to a conventional loss function that does not support the ability to separately resolve these losses as further described in the corresponding section in the following discussion.
[0048] Generally, the functionality, features, and concepts described with respect to the examples above and below can be adopted in the context of the example processes described in this section. Further, the functionality, features, and concepts described with respect to the different figures and examples in this document can be interchanged with each other and are not limited to implementation in the context of a particular figure or process. Additionally, the blocks associated with the different representative processes and corresponding figures in this document can be applied and / or combined together in different ways. Accordingly, the individual functionality, features, and concepts described with respect to the different example environments, devices, components, figures, and processes in this document can be used in any suitable combination, and are not limited to the particular combinations represented by the enumerated examples in this specification.
[0049] Query-Based Training Dataset
[0050] Figure 2 Depicts system 200 in an example implementation, where service provider system 102 generates a query-based training dataset 202 based on a text query and a digital image associated with the text query. Figure 3 Depicts process 300 in an example implementation, where the training dataset is used to train a model using machine learning. The training dataset is generated based on digital images and text queries that are used to locate the digital images as part of a search.
[0051] The following discussion describes techniques that can be implemented using the previously described systems and devices. Aspects of the process can be implemented in hardware, firmware, software, or a combination thereof. The process is shown as a set of blocks specifying operations to be performed by one or more devices, and is not necessarily limited to the order shown for performing the operations corresponding to the respective blocks. In the various parts of the following discussion, reference will be made interchangeably to Figures 1 to 3 .
[0052] The accuracy of training data is one of the driving factors when using machine learning to train a model 120 to perform functions accurately. Another driving factor is to obtain a sufficient quantity of accurate training data. However, in practice, this can be difficult. For example, conventional techniques for labeling digital images typically rely on users manually indicating which objects are included in the digital image, the characteristics of the objects, the emotions evoked by the objects, etc. However, this can vary from person to person, and it is also expensive when attempting to perform machine learning against a sufficient quantity of digital images and associated text.
[0053] Therefore, Figure 2The illustrated system 200 is configured to generate a query-based training dataset 202 based on digital image search. In the illustrated example, the service provider system 102 includes a search module 114 that is configured to support searching for digital images 110 from a storage device 112, e.g., local to or remote from the service provider system 102. For example, the service provider system 102 may support a storage device 112 that includes digital images 110 as an “inventory” and provide free access to the “inventory,” paid purchase (e.g., subscription or “cataloging”) of the “inventory,” etc. In another example, the service provider system 102 implements the search module 114 portion of a search engine system that locates digital images maintained by a third-party system. Other implementation examples are also contemplated.
[0054] In the illustrated example, a text query 116—“running shoes”—is input via a user interface 118 of the client device 104 as previously described with respect to Figure 1 As a response, the client device 104 receives search results from the service provider system 102 that include digital images 124(1) through 124(4) as displayed in the user interface 118. A user input is then received via the user interface 118, which is illustrated as a tap gesture detected via touchscreen functionality of the client device 104 that selects digital image 124(3). The training data generation module 126 uses the user input to determine the association of digital image 124(3) with the text of the text query 116. Accordingly, the training data generation module 126 can use this correspondence to generate a query-based training dataset 202 based on first data 206 and second data 208, where the data 206 describes the text query 116 and the data 208 describes the selected digital image 124(3) in the search results. In this way, the training data generation module 126 can obtain a large number of digital images associated with a large number of different texts and, likewise, overcome the limitations of conventional training data.
[0055] For example, the training data generation module 126 can be as described in Figure 3Receive a plurality of text queries (block 302) used to initiate multiple digital image searches from a large number of client devices 104 as illustrated. The training data generation module 126 also receives a plurality of digital images selected by the user from the search results generated by the multiple digital image searches (e.g., via gestures, cursor control devices, spoken words) (block 304). Thus, the training data generation module 126 receives a plurality of digital images and text queries respectively associated with the plurality of digital images. In this way, the digital images and text can cover a wide range of the following digital image and text associations: It is difficult, if not impossible, to obtain these digital image and text associations using conventional manual labeling methods and even conventional automated techniques that can support limited text instances.
[0056] The training data generation module 126 generates a training data set 128 based on the plurality of text queries and the plurality of digital images (block 306). For example, the plurality of digital images can be regarded as positive digital image samples of the associated text queries. The training data generation module 126 can also generate negative digital image samples to be used as part of the training, and further discussion on this can be found in the negative digital image sample generation section in the following discussion. In the illustrated example, this results in a query-based training data set 202.
[0057] As illustrated in Figure 1 the training data set 128 is passed from the training data generation module 126 to the machine learning training module 130. The machine learning training module 130 is configured to use the machine learning training model 120 based on the loss function 132 by using the training data set (block 308). Once the model 120 is trained, the search module 114 can then use the model 120 to generate subsequent search results (block 310), e.g., in response to subsequent search queries.
[0058] The training data generation module 126 can also be employed to generate a training data set 128 using other sources of digital images and associated text. For example, a query-based training data set 202 can include "noise" caused by selecting digital images that do not accurately reflect the text in the text query. This can be caused by a user being interested in search results of digital images that do not exactly correspond to the text query. For example, a user can enter the text query "running shoes", but receive as part of the search results a digital image 124(1) of a running dog of a breed the user is interested in. Thus, the user's selection of the digital image 124(1) does not accurately reflect the association between the text query and the digital image, but rather indicates the user's interest in the image. Thus, the data describing the association between the text query 116 and the digital image 124(1) may introduce "noise". In other instances, for a text query containing a large amount of text, the search module 114 may not return results, for example, as may occur in some stock digital image systems.
[0059] Accordingly, the training data generation module 126 can also obtain digital images associated with text that can be used to supplement the training data set 128. One such example includes digital images having associated captions (e.g., title bars) that the training data generation module 126 uses to generate a caption-based training data set 422. In fact, the caption associated with a digital image can include a large amount of text that is used to describe the object, object characteristics, location, induced mood, etc. in the digital image. By including the caption-based training data set 422 with the query-based training data set 202, the training data set 128 can address the noise introduced in the query-based training data set 202, support the use of "long sentences", address text sequences, and thus be able to understand text queries with improved accuracy and efficiency, for example, be able to support "girl cat" and "girl holding a cat" as text queries. A further discussion of the generation of the training data set 128 is included in the next section.
[0060] Negative Digital Image Sample Generation
[0061] Figure 4 Depicts system 400 in an example implementation, where a training data set is generated that includes negative digital image samples selected based on positive digital image samples and associated text. Figure 5 Depicts process 500 in an example implementation, where negative digital image samples are generated based on a comparison of text associated with the negative digital image samples and text associated with the positive digital image samples.
[0062] The following discussion describes techniques that can be implemented using the previously described systems and devices. Aspects of the process can be implemented in hardware, firmware, software, or a combination thereof. The process is shown as a set of blocks that specify operations to be performed by one or more devices, and is not necessarily limited to the order of operations shown for performing the corresponding blocks. In various parts of the following discussion, reference will be made interchangeably to Figure 1 , Figure 4 and Figure 5 .
[0063] When training the model 120 via the machine learning training module 130, positive digital image samples and negative digital image samples are used as part of a triplet loss to adjust the weights of the neurons in the neural network of the model 120. This is done to ensure that, for the embedding space implemented by the model 120, examples with the same or similar text (i.e., digital images) are closely clustered together in the embedding space (i.e., the single unified text and digital image embedding space 122), while examples with different text are not close together in the embedding space, and such that tighter clusters are formed.
[0064] In this section, techniques for generating negative digital image samples that improve accuracy and computational efficiency when training the model 120 via the machine learning training module 130 are described. The training data generation module 126 accomplishes this automatically and without user intervention by: generating negative digital image samples having semantic and / or visual meanings similar to, but not exactly the same as, those of the positive digital image samples, and thus, improving the ability of the model 120 to discriminate between these samples.
[0065] Beginning, the training data generation module 126 receives a plurality of digital images and associated text 402 (block 502). The plurality of digital images and associated text 402 can include digital images and text queries 404, digital images and captions 406, and other examples of digital image and text associations that can be used to generate the training dataset 128.
[0066] Then, the training data generation module 126 automatically and without user intervention generates the training dataset 128 based on the plurality of digital images and associated text 402 (block 504). First, the positive digital image generation module 408 selects positive digital image samples 410 from the plurality of digital images (block 506). This can be done by selecting any digital image from the plurality of digital images and associated text 402 using a queue or the like.
[0067] Then, the negative sample generation module 412 generates negative digital image samples 414 (block 508) from the multiple digital images and associated text 402 based on the positive digital image samples 410. The negative sample generation module 412 can perform this in various ways, examples of which include filtering the multiple digital images.
[0068] In one example of the filtering, a subset of the multiple digital images is generated by the negative sample generation module 412. This is performed by removing from the multiple digital images those digital images having at least one text item that does not include stop words and is also included in the text associated with the positive digital image samples (block 510), and the remaining digital images form the subset. For example, if the text associated with the positive digital image samples is "a person on a motorcycle", removing the stop words "on" and "a" results in the text items "person" and "motorcycle", i.e., the "center points". Thus, each digital image in the multiple digital images associated with text including "person" or "motorcycle" is removed to form the subset, i.e., other images are filtered out from the multiple digital images.
[0069] Then, the negative sample generation module 412 selects negative digital image samples 414 from the subset (block 514). For example, the negative sample generation module 412 can select "N" negative samples based on the least square distance from the positive digital image samples 410 using the corresponding text embeddings, which are generated using a convolutional neural network. This is an example of "hard" negative image selection, which in implementation is used to generate a query-based training dataset 202 from the digital images and text queries 404, and the query-based training dataset 202 is used as part of multi-task training as further described below.
[0070] In another example, the negative sample generation module 412 generates a subset of the plurality of digital images that does not include digital images having each text item, where the text items do not include stop words and are also included in the text associated with the positive digital image samples 410 (block 512). In other words, digital images that do have each text item are filtered out from the plurality of digital images, and the remaining digital images form a subset. Negative digital image samples 414 are then selected from the subset (block 514). Continuing with the previous example, if the text associated with the positive digital image samples is "a person on a motorcycle", removing the stop words "on" and "a" yields the text items "person" and "motorcycle", i.e., the "center points". Then a subset is generated from the plurality of digital images and the associated text 402, and the remaining digital images and associated text 402 are not associated with text including "person" and "motorcycle". This is considered to generate even "harder" negative samples, and in an implementation, is used for the digital images and captions 406 to generate a caption-based training dataset 422 as part of multi-task training. For example, this can be used to address a technical challenge because the amount of text typically observed for captions is greater than the amount of text typically observed for text queries, and thus, for captions with improved robustness, this generates negative digital image samples 414.
[0071] In this example, the negative sample generation module 412 further selects negative digital image samples 414 from the subset (block 514). For example, the negative sample generation module 412 can select "N" negative samples based on the least squared distance to the positive digital image samples 410 using the corresponding image embeddings, where the corresponding image embeddings are generated using a convolutional neural network (CNN).
[0072] The triplet formation module 416 generates triplets such as including the positive digital image samples 410, the text 420 in the plurality of texts associated with the positive digital image samples 410, and the negative digital image samples 414 (block 516). For example, the text extraction module 418 can extract the text 420 from the plurality of digital images and the associated text 402 corresponding to the positive digital image samples 410. In this way, the training data generation module 126 generates a training dataset 128 from the plurality of digital images and the associated text 402, and the training dataset 128 can include a query-based training dataset 202 and a caption-based training dataset 422. The query-based training dataset 202 and the caption-based training dataset 422 can be used to train the model 120 using a loss function 132 as part of machine learning as further described in the next section (block 518).
[0073] Loss Function
[0074] Figure 6Depicts system 600 in an example implementation, which shows a machine learning training module 130 performing multi-task training on a model 120 using multi-task training based on the following training datasets: a query-based training dataset 202 and a caption-based training dataset 422. Figure 7 Depicts system 700, which more particularly shows the operation of machine learning training module 130 in generating embeddings and using a loss function 132. Figure 8 Depicts in more detail Figure 7 System 800 that more particularly shows the following operation of the illustrated text embedding module: generating text embeddings from text associated with positive digital image samples. Figure 9 Depicts a comparison between a conventional triplet loss and a positive-aware triplet ranking loss as described herein. Figure 10 Depicts process 1000 in an example implementation, where a model is trained based on a loss function that separately addresses the loss between text embeddings and positive image embeddings from the loss between text embeddings and negative image embeddings.
[0075] The following discussion describes techniques that can be implemented using the previously described systems and devices. Each process can be implemented in hardware, firmware, software, or a combination thereof. The processes are shown as a set of blocks specifying operations to be performed by one or more devices and are not necessarily limited to the order shown for performing the operations of the corresponding blocks. In various parts of the following discussion, reference will be made interchangeably to Figure 1 and Figures 6 to 10 .
[0076] As previously described, a multi-task training approach can be taken when training model 120 by machine learning training module 130. This is performed in Figure 6 by using a training dataset 128 that includes a query-based training dataset 202 and a caption-based training dataset 422. Each of these datasets includes corresponding triplets 602, 604 of positive digital image samples, text associated with the corresponding positive digital image samples, and negative digital image samples as described in the previous section. In this way, machine learning training module 130 is configured to capture user intent regarding the association of text queries with corresponding digital images from the query-based training dataset 202 and also use the caption-based training dataset 422 to create embeddings of long text sequences (e.g., sentences). Thus, once model 120 is trained, model 120 is able to resolve text and text sequences with improved efficiency and accuracy, e.g., able to resolve the difference between "girl cat" and "girl holding a cat".
[0077] As part of this, the machine learning training module 130 generates a single unified text and digital image embedding space 122 into which the digital images and the associated text are projected together. For example, the machine learning training module 130 can utilize pre-trained architectures and train these pre-trained architectures on a large corpus of digital images to make predictions on labels, examples of which include VGG-19, ResNet-152, ResNet-50, etc. For example, the layer located before the last activation layer in these architectures (i.e., the SoftMax layer) can be utilized by the machine learning training module 130 as a common image-based embedding space. For this purpose, a modified version of the triplet loss is used as the loss function 132 to train the model 120.
[0078] Figure 7 FIG. depicts a system 700 that more particularly illustrates the operation of the machine learning training module 130 in generating embeddings and using the loss function 132 to train the model 120. Continuing with the previous example, the machine learning training module 130 is configured to perform multi-task training in which samples are separately taken from a query-based training data set 202 and a caption-based training data set 422. For example, the samples can form triplets that include a positive digital image sample, the text associated with the positive digital image sample, and a negative digital image sample generated based on the positive digital image sample.
[0079] Thus, as described above, the service provider system 102 can separately receive a plurality of digital images and a plurality of texts associated with the plurality of digital images (block 1002), e.g., text queries, captions, etc. Then, the training data generation module 126 is utilized to generate a training data set 128 based on the plurality of digital images and the plurality of texts. The training data set 128 includes positive digital image samples, the text associated with the positive digital image samples among the plurality of texts, and negative digital image samples (block 1004). Then, the training data generation module 126 outputs the training data set 128, and the machine learning training module 130 receives the training data set 128 as input.
[0080] The machine learning training module 130 uses the training data set 128 to use the machine learning training model 120 based on the loss function 132 (block 1006). The machine learning training module 130 begins training the model 120 by forming embeddings (e.g., vectors) of text and digital images using the text encoder 702 and the digital image encoder 704 to generate a text embedding 706 and positive image embeddings 708 and negative image embeddings 710, respectively (block 1008). In the illustrated example, the text encoder 702 uses a recurrent neural network (RNN) language encoder 712 to generate the text embedding 706 (e.g., a vector of length 2048) based on the text. An RNN is a type of neural network where the connections between nodes are for a directed graph along a time series and can use internal states to process a sequence of inputs. In this way, the text embedding 706 can capture the order of the text (e.g., within a text query or text input), which is not possible in a label-based approach.
[0081] Figure 8 System 800 is depicted that more particularly illustrates an example of the operation of the text encoder 702. The text encoder 702 includes a pre-trained word embedding module 802 that has a dictionary that includes embeddings for text within a particular language, an example of an embedding for text being referred to as "Fasttext". The word embeddings generated by the pre-trained word embedding module 802 provide semantic information about the text to the model.
[0082] Then, the output of the pre-trained word embedding module 802 is provided to a set of stacked long short-term memory (LSTM) cells 804 to capture the sequential information of the text 816 with respect to each other. The output of the last cell in the stacked LSTM cells 804 is output to a fully connected layer 806 to transform the vector size (e.g., from 300 to 2048), which results in the text embedding 706. The machine learning training module 130 can utilize this to generate text embeddings for text queries in the query-based training data set 202, captions in the caption-based training data set 422, etc.
[0083] In Figure 7 the illustrated example, the digital image encoder 704 is configured to use a convolutional neural network (CNN) image encoder 714 to generate the positive image embeddings 708 and negative image embeddings 710 (e.g., vectors). The CNN image encoder 714 includes a series of pre-trained convolutional layers that have filters and pooling layers for extracting and learning features of the digital images to generate embeddings in the image embedding space. As a result, the text embedding 706 and the positive image embeddings 708, negative image embeddings 810 can be directly used as part of a single unified text and digital image embedding space 122 implemented by the model 120.
[0084] Once the text embedding 706 and the positive image embedding 708, negative image embedding 710 are generated, the model 120 is trained using the loss function 132. The loss function 132 of the machine learning training module 130 is configured to determine the loss between the text embedding and the positive image embedding separately from the loss between the text embedding and the negative image embedding (block 1010).
[0085] For example, as in Figure 7 The loss function 132 illustrated in includes an L2 716 loss (e.g., squared distance) that is used to determine the loss between the text embedding 706 and the positive image embedding 708 separately from the L2 718 (e.g., squared distance) loss determined between the text embedding 706 and the negative image embedding 710. In the current discussion, this is referred to as the “positive-aware triplet ranking loss”, which can be expressed as follows:
[0086] Positive perception triple ranking loss = s p +max(0,margin–s n )
[0087] where the squared distance between the positive image embedding 708 and the text embedding 706 is “s p ”, and the squared distance between the negative image embedding 710 and the text embedding 706 is “s n ”.
[0088] The conventional triplet loss function is configured by increasing “s p ” and “s n "The values of the two make "s p –s n " is minimized. Therefore, in a conventional triplet loss function, when two values increase, the difference automatically increases. However, in the positive-aware triplet ranking loss illustrated as the loss function 132 described herein, the loss is solved separately. Therefore, the positive-aware triplet ranking loss is configured to minimize "s n ” (i.e., the loss between the text embedding 706 and the negative text embedding 706) to maximize the “s p ” (i.e., the loss between the text embedding 706 and the positive image embedding 708) is minimized. This allows the positive image embedding 708 to improve its similarity to the text embedding 706, e.g., to be in the same cluster, and at the same time n " is maximized (i.e., increasing the dissimilarity with the negative image embedding) to make the clusters tighter.
[0089] In instances where multiple negative samples are used, the machine learning training module 130 selects the top "N" samples with the smallest squared distance that have not been rejected (e.g., filtered), rather than the top samples. Then, the loss function 132 can be expressed as:
[0090] Positive aware triplet ranking loss = s p + ∑i(max(0, margin – S ni ))
[0091] In an implementation, the number of samples is increased by a defined number at defined time points (such as every ten epochs).
[0092] Figure 9 Graphical example 900 is depicted, which contrasts the losses calculated using the conventional triplet loss function 902 and the positive aware triplet loss function 904 as described above. As shown, for the conventional triplet loss function 902, the difference between the negative loss 906 and the positive loss 908 tracks each other. This is because the conventional triplet loss function is configured to minimize "s p " and "s n " by increasing the values of "s p – s n ". Thus, in the conventional triplet loss function, when both loss values increase, the difference tracks these increases.
[0093] However, when minimizing the positive loss and maximizing the negative loss, the difference between the negative loss 906 and the positive loss 908 for the positive aware triplet loss function 904 increases over time. In other words, during training, the distance between the positive image embedding 708 and the text embedding 706 decreases over time, while the distance between the negative image embedding 710 and the text embedding 706 increases. In instances of multi-task training, the post-aware triplet loss function can employ different margins for the losses for the query-based training dataset 202 and the caption-based training dataset 422. Then, the L2 losses 716, 718 are averaged into Figure 7 the loss 720 in, and it is backpropagated (722) through the network, e.g., to train the text encoder 702 to utilize the image embedding space of the digital image encoder 704. Once trained, the model 120 can be used to support various functionalities, such as to generate search results (box 1012), digital image retrieval, machine translation, text retrieval, speech recognition, text summarization, etc., such that visual intuition is supported in these technologies to address the visually expressed "content" in text.
[0094] Thus, as described above, text-to-visual machine learning embedding techniques are configured to overcome challenges in conventional techniques in various ways. These techniques include: using query-based training data, which can expand the availability and type of training data available for training a model. The use of query-based training data can also be extended by using caption-based training data as part of multi-task learning, which improves training accuracy by limiting noise in the query-based training data and supports the use of long text sequences.
[0095] The generation of negative digital image samples is also described, which improves accuracy when using machine learning to train a model by having a similar semantic and / or visual meaning as positive digital image samples, but not having exactly the same components as the positive digital image samples. This training data can then be used to generate a model that supports a single unified text and digital image embedding space, which is configured to treat text and digital images as the same entity and thus overcomes the limitations of conventional techniques that are based solely on text.
[0096] A loss function that also supports improved accuracy and computational efficiency is also described by separately processing the loss computed between a positive image embedding generated from a positive digital image sample and the text embedding from the negative image embedding generated from the negative digital image sample and the following text embedding: the text embedding is computed based on the text associated with the positive digital image sample. This allows the distance between the positive image embedding and the text embedding to decrease over time, while the distance between the negative image embedding and the text embedding increases, thereby improving model accuracy.
[0097] Example Systems and Devices
[0098] Figure 11 An example system including an example computing device 1102 is generally illustrated at 1100, and the example computing device 1002 represents one or more computing systems and / or devices that can implement the various techniques described herein. This is illustrated by including a training data generation module 126, a machine learning training module 130, and a model 120. The computing device 1102 can be, for example, a server of a service provider, a device associated with a client (e.g., a client device), a system-on-chip, and / or any other suitable computing device or computing system. Further, the computing device 1102 can implement a platform 1116 and resources.
[0099] As shown in the figure, the exemplary computing device 1102 includes a processing system 1104, one or more computer-readable media 1106, and one or more I / O interfaces 1108 that are communicatively coupled to each other. Although not shown, the computing device 1102 may further include a system bus or other data and command transmission systems that couple the various components to each other. The system bus may include any bus structure or combination of bus structures in different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus using any bus architecture in various bus architectures. Various other examples are also contemplated, such as control lines and data lines.
[0100] The processing system 1104 represents functionality for performing one or more operations using hardware. Thus, the processing system 1104 is illustrated as including hardware elements 1110, which may be configured as processors, functional blocks, etc. This may include implementation as an application specific integrated circuit in hardware or other logic devices formed using one or more semiconductors. The hardware elements 1110 are not limited by the materials from which they are formed or the processing mechanism in which they are deployed. For example, a processor may be composed of (multiple) semiconductors and / or transistors (e.g., electronic integrated circuits (ICs)). In this context, processor-executable instructions may be electronically executable instructions.
[0101] The computer-readable media 1106 is illustrated as including a memory / storage device 1112. The memory / storage device 1112 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage device 1112 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic disks, etc.). The memory / storage device 1112 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). The computer-readable media 1106 may be configured in various other ways as further described below.
[0102] (Multiple) input / output interfaces 1008 represent functionality for allowing a user to enter commands and information into computing device 1202 and also for allowing information to be presented to the user and / or other components or devices using a variety of input / output devices. Examples of input devices include: keyboards, cursor control devices (e.g., mice), microphones, scanners, touch functionality (e.g., capacitive sensors or other sensors configured to detect physical touch), cameras (e.g., the camera may use visible or invisible wavelengths (such as infrared frequencies) to recognize movement as a gesture not involving touch), and the like. Examples of output devices include: display devices (e.g., monitors or projectors), speakers, printers, network cards, haptic response devices, and the like. Thus, computing device 1102 can be configured to support user interaction in a variety of ways as further described below.
[0103] Various techniques may be described herein in the general context of software, hardware elements, or program modules. Generally, such modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. As used herein, the terms "module," "functionality," and "component" generally refer to software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that these techniques can be implemented on a variety of commercial computing platforms having a variety of processors.
[0104] Implementations of the described modules and techniques may be stored on some form of computer-readable medium or may be transmitted on some form of computer-readable medium. Computer-readable media may include various media that can be accessed by computing device 1102. By way of example and not limitation, computer-readable media may include "computer-readable storage media" and "computer-readable signal media."
[0105] In contrast to merely signal transmission, carrier, or the signal itself, a "computer-readable storage medium" can refer to a medium and / or device capable of storing information permanently and / or non-transiently. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware (such as volatile and non-volatile, removable and non-removable media) and / or storage devices implemented in accordance with a method or technology suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media can include, but are not limited to: RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile discs (DVDs), or other optical storage devices, hard disks, magnetic tape cartridges, tapes, magnetic disk storage devices, or other magnetic storage devices, or other storage devices suitable for storing the desired information and accessible by a computer, tangible media, or articles.
[0106] A "computer-readable signal medium" can refer to a signal-bearing medium configured to transmit instructions (such as via a network) to the hardware of computing device 1102. A signal medium can generally embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, data signal, or other transport mechanism. A signal medium also includes any information delivery medium. The term "modulated data signal" refers to a signal having one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic wireless media, RF, infrared wireless media, and other wireless media).
[0107] As previously described, hardware elements 1110 and computer-readable medium 1106 represent modules, programmable device logic, and / or fixed device logic implemented in hardware form (such as to execute one or more instructions) that can be employed in some embodiments to implement at least some aspects of the techniques described herein. Hardware can include components of an integrated circuit or system-on-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implementations in silicon or other hardware. In this context, hardware can operate as a processing device to execute program tasks defined by instructions and / or logic implemented in hardware and hardware utilized to store instructions for execution, e.g., the computer-readable storage medium previously described.
[0108] The various techniques described herein can also be implemented using combinations of the foregoing. Accordingly, software, hardware, or executable modules can be implemented as one or more instructions and / or logic implemented on a computer-readable storage medium of some form and / or via one or more hardware elements 1110. Computing device 1102 can be configured to implement specific instructions and / or functions corresponding to software modules and / or hardware modules. Accordingly, the implementation as a software-executable module by computing device 1102 can be implemented at least in part in hardware, e.g., by using a computer-readable storage medium of processing system 1104 and / or hardware elements 1110. Instructions and / or functions can be executed / operated by one or more articles of manufacture (e.g., one or more computing devices 1102 and / or processing systems 1104) to implement the techniques, modules, and examples described herein.
[0109] The techniques described herein can be supported by various configurations of computing device 1102 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part by using a distributed system, such as via platform 1116 on "cloud" 1114 as described below.
[0110] Cloud 1114 includes and / or represents platform 1116 for resources 1118. This platform 1116 abstracts the underlying functionality of the hardware resources (e.g., servers) and software resources of cloud 1114. Resources 1118 can include applications and / or data that can be utilized when computer processing is executed on servers remote from computing device 1102. Resources 1118 can also include services provided over the Internet and / or over a subscriber network such as a cellular or Wi-Fi network.
[0111] Platform 1116 can abstract resources and functionality to connect computing device 1102 with other computing devices. Platform 1116 can also be used to abstract the scaling of resources to provide corresponding scaling levels for meeting the demand for resources 1118 implemented via platform 1116. Accordingly, in an interconnected device embodiment, the implementation of the functionality described herein can be distributed across system 1100. For example, the functionality can be implemented in part on computing device 1102 and via platform 1116 that abstracts the functionality of cloud 1114.
[0112] Conclusion
[0113] Although the invention has been described in language specific to structural features and / or methodological acts, it is to be understood that the invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed invention.
Claims
1. A method implemented by a computing device in a digital media machine learning model training environment, the method comprising: Receiving, by the computing device, a plurality of text queries that are used to initiate multiple digital image searches; Generating, by the computing device, a plurality of filtered text queries by filtering stop words from the plurality of text queries; Generating, by the computing device, a training data set based on the plurality of filtered text queries and a plurality of digital images generated by the multiple digital image searches, the training data set including: Positive digital image samples located using a first respective one of the filtered text queries; and Negative digital image samples using a second respective one of the filtered text queries, the second respective one of the filtered text queries sharing at least one text item with the first respective one of the filtered text queries and not sharing at least one other text item with the first respective one of the filtered text queries; Using, by the computing device, the training data set to train a machine learning model based on a loss function; and Using, by the computing device, the model to generate subsequent search results.
2. The method according to claim 1, wherein the training of the model is based on the plurality of filtered text queries and the plurality of digital images to generate a single unified text and digital image embedding space.
3. The method according to claim 1, wherein the generation of the training data set includes: Selecting positive digital image samples from the plurality of digital images; And Generating negative digital image samples from the plurality of digital images based on the positive digital image samples.
4. The method according to claim 3, wherein the stop words are common words that are irrelevant to the performance of the digital image search.
5. The method according to claim 1, wherein the generation of the training data set comprises: Generate a title-based training data set that has titles associated with corresponding multiple digital images.
6. The method according to claim 5, wherein the generation of the title-based training data set includes: Selecting the positive digital image samples from the corresponding multiple digital images; And Generating the negative digital image samples from the corresponding multiple digital images based on the positive digital image samples.
7. The method according to claim 6, wherein the generation of the negative digital image samples includes: Generating a filtered title by filtering stop words from the title; Generating a subset of the corresponding multiple digital images by excluding digital images from the corresponding multiple digital images that have each text item included in the filtered title associated with the positive digital image sample; And Selecting the negative digital image samples from the subset.
8. The method according to claim 1, wherein the training comprises: Generating positive image embeddings from the positive digital image samples, generating text embeddings from the first respective one of the filtered text queries associated with the positive digital image samples, and generating negative image embeddings from the negative digital image samples.
9. The method according to claim 8, wherein the loss function is a triplet loss function, and the triplet loss function separately addresses the loss between the text embedding and the positive image embedding from the loss between the text embedding and the negative image embedding.
10. A system in a digital media machine learning model training environment, comprising: a search module at least partially implemented in hardware to generate a plurality of digital images generated by multiple digital image searches; a training data generation module at least partially implemented in hardware to generate a training data set, the training data generation module comprising: a positive sample generation module configured to select positive digital image samples from the plurality of digital images; and a negative sample generation module configured to: generate filtered text by filtering the text associated with the plurality of digital images; generate a subset of the plurality of digital images, the subset including digital images from the plurality of digital images associated with at least one filtered text item, the at least one filtered text item also being included in the filtered text associated with the positive digital image samples and not sharing at least one other filtered text item with the positive digital images; and select negative digital image samples from the subset; a machine learning training module at least partially implemented in hardware to train a model as part of machine learning based on the training data set using a loss function.
11. The system according to claim 10, wherein the filtered text describes a text query used to locate a corresponding digital image in the plurality of digital images as part of a search.
12. The system according to claim 10, wherein the filtered text describes a title associated with the respective digital image.
13. The system according to claim 10, wherein the machine learning training module is configured to: generate a positive image embedding from the positive digital image samples, generate a text embedding from the filtered text associated with the positive digital image samples, and generate a negative image embedding from the negative digital image samples.
14. The system according to claim 13, wherein the loss function is a triplet loss function, and the triplet loss function separately processes the loss between the text embedding and the positive image embedding from the loss between the text embedding and the negative image embedding.
15. A method implemented by a computing device in a digital media machine learning model training environment, the method comprising: receiving, by the computing device, a plurality of digital images and a plurality of texts associated with the plurality of digital images, respectively; generating, by the computing device, a training data set based on the plurality of digital images and the plurality of texts, the training data set having a first training data set of queries and a second training data set based on titles, and including positive digital image samples, the texts in the plurality of texts associated with the positive digital image samples, and negative digital image samples; The computing device uses the training dataset to train a machine learning model based on a loss function, the training including: generating a text embedding from the text, a positive image embedding from the positive digital image samples, and a negative image embedding from the negative digital image samples; and using the loss function to separately determine a loss between the text embedding and the positive image embedding for the first training dataset from a loss between the text embedding and the negative image embedding of the second training dataset.
16. The method according to claim 15, wherein the training trains the model based on the plurality of texts and the plurality of digital images to achieve a single unified text and digital image embedding space.
17. The method according to claim 15, wherein during the training, a distance of the loss between the text embedding and the positive image embedding decreases, and a distance of the loss between the text embedding and the negative image embedding increases.
18. The method according to claim 15, wherein: the first training dataset is a query-based training dataset, the query-based training dataset includes a plurality of text queries and a plurality of digital images, the plurality of text queries are used to initiate multiple digital image searches, and the plurality of digital images are selected by a user from search results generated by the multiple digital image searches; and the second training dataset is a caption-based training dataset, the caption-based training dataset includes a corresponding plurality of digital images and captions associated with the corresponding plurality of digital images.
19. The method according to claim 18, wherein the loss function calculates a loss for the query-based training dataset separately from a loss for the caption-based training dataset.
20. The method according to claim 19, wherein a loss for the training dataset is calculated by averaging the loss for the query-based training dataset and the loss for the caption-based training dataset.
Citation Information
Patent Citations
Modeling Semantic Concepts in an Embedding Space as Distributions
US20170206465A1
Method And Apparatus Of Establishing Image Search Relevance Prediction Model, And Image Search Method And Apparatus
US20170330054A1