Multi-dimensional digital content search

Through multi-dimensional digital content search technology, using the position and coordinate weights in the multi-dimensional continuous space, the problem of difficulty in positioning multiple emotional and characteristic digital content in the prior art is solved, and more efficient and accurate search results are achieved.

CN113836382BActive Publication Date: 2025-07-11ADOBE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110377933.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-24
Filing Date
2021-04-08
Publication Date
2025-07-11
Estimated Expiration
2041-04-08

AI Technical Summary

Technical Problem

Existing search technologies are difficult to accurately locate digital content items with a variety of difficult-to-express concepts, resulting in inefficient use of computing and network resources.

Method used

The multi-dimensional digital content search technology is adopted to specify the weight of the search criteria by defining the positions and coordinates in a multi-dimensional continuous space, allowing users to perform complex emotional and multi-standard searches with a single input.

Benefits of technology

Improves the accuracy and efficiency of searches, and better positioning digital content items with multiple emotional and characteristics, overcoming consistency challenges in conventional technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113836382B_ABST
    Figure CN113836382B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to multi-dimensional digital content search. Multi-dimensional digital content search techniques are described that support the ability of a computing device to perform searches with increased granularity and flexibility compared to conventional techniques. In one example, a control defining a multi-dimensional (e.g., two-dimensional) continuous space is implemented by a computing device. Positions in the multi-dimensional continuous space are available for different search criteria by applying different weights to criteria associated with axes. Thus, user interaction with this control can be used to define positions and corresponding coordinates that can act as weights for search criteria in order to perform a search of digital content using a single user input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to multi-dimensional digital content search. Background Art

[0002] Search is one of the main techniques used by computing devices to locate a specific digital content item from thousands or even tens of millions of digital content instances. For example, a computing device can use search to locate a digital image from millions of stock digital images, a digital music item from a song library, a digital movie from thousands of movies available on an online streaming service, and so on. As a result, digital search can be implemented to solve situations involving multiple digital content items in a way that humans do not actually perform.

[0003] However, search implemented by computing devices faces numerous challenges, one of which involves the ability to determine the user's intent in a search query and locate digital content that matches that intent. For example, conventional search techniques typically rely on the ability to match text received in a search query with text associated with digital content. Although this technique can well locate digital content with a specific object (e.g., for the search query "dog"), it may fail when encountering concepts that are not easily expressed in text, such as emotions, relative quantities of search criteria, etc. Therefore, conventional search techniques are usually inaccurate and result in inefficient use of computing and network resources due to repeated attempts to locate a specific digital content item of interest when faced with these concepts. Summary of the Invention

[0004] Multi-dimensional digital content search techniques are described, which support the ability of a computing device to perform searches with higher granularity and flexibility than conventional techniques. In one example, a control is implemented by a computing device that defines a multi-dimensional (e.g., two-dimensional) continuous space. Positions in the multi-dimensional continuous space can be used to specify weights applied to search criteria associated with axes. Thus, user interaction with the control can be used to define positions and corresponding coordinates, which can serve as weights for search criteria in order to perform a search of digital content using a single user input.

[0005] This "Summary of the Invention" introduces some concepts in a simplified form, which will be further described in the following "Detailed Description". Thus, this "Summary of the Invention" is neither intended to identify the essential features of the claimed subject matter nor to be used to help determine the scope of the claimed subject matter. Brief Description of the Drawings

[0006] The detailed description is described with reference to the accompanying drawings. Entities represented in the drawings may indicate one or more entities, and thus entities in the singular or plural form may be referred to interchangeably in the discussion.

[0007] Figure 1 is an illustration of a digital media search environment in an example implementation operable to employ digital content search techniques;

[0008] Figure 2 depicts an example of controls configured to support a multi-dimensional continuous space for searching using emotion; Figure 1 of;

[0009] Figure 3 depicts a system in an example implementation that more particularly illustrates the operation of a search I / O module and a digital content search system when performing a multi-dimensional digital content search; Figure 1 of;

[0010] Figure 4 depicts an example of a multi-dimensional digital content search involving emotion;

[0011] Figure 5 depicts another example of a multi-dimensional digital content search involving emotion;

[0012] Figure 6 is a flow chart depicting a process in an example implementation where controls including a representation of a multi-dimensional continuous space are utilized as part of a digital content search;

[0013] Figure 7 The Figure 3 machine learning model is more particularly depicted as an ensemble model including an image model and a tag-based model;

[0014] Figure 8 depicts an example of emotion tag coordinates defined with respect to a pleasure X-axis and an excitement Y-axis;

[0015] Figure 9 depicts an example of tags associated with a digital image;

[0016] Figure 10 depicts another example of tags associated with a digital image; and

[0017] Figure 11 shows an example system including various components of an example device that may be implemented as any type of computing device described and / or utilized with reference to Figures 1 to 10 to implement embodiments of the techniques described herein. DETAILED DESCRIPTION

[0018] Overview

[0019] Searches implemented by a computing device can be used to locate a specific digital content item from millions of examples in real time. Thus, the search implemented by the computing device supports the ability of a user to interact with the digital content, which would otherwise not be possible, i.e., not performable by humans alone. However, conventional search techniques implemented by a computing device often fail when encountering concepts that are difficult to express (e.g., in text).

[0020] For example, a computing device can use a text search query "dog" to locate numerous examples of digital images associated with the tag "dog". Similarly, a search for a single emotion and the recognition of an object (such as "happy dog") can return digital images with both the tags "dog" and "happy". However, conventional techniques do not support the ability to specify weights for search criteria, nor weights that are applied together to multiple search criteria. For example, search queries including "happy enthusiastic dog" or "sad calm girl" typically fail with conventional search techniques because they cannot address multiple emotions simultaneously and result in inefficient use of network and computing resources.

[0021] Accordingly, a multi-dimensional digital content search technique is described that supports the ability of a computing device to perform searches with a higher granularity and flexibility than conventional techniques. In one example, a control is implemented by the computing device that defines a continuous space involving at least two search criteria. The first and second axes of the control can, for example, correspond to positive and negative amounts of excitement emotion and pleasure emotion, respectively.

[0022] In this way, the control defines a multi-dimensional (e.g., two-dimensional) continuous space. A position in the multi-dimensional continuous space can be used to specify weights applied to search criteria associated with the axes. Continuing with the emotion example above, emotions such as happy, elated, excited, nervous, angry, disappointed, sad, bored, tired, calm, relaxed, and satisfied (i.e., content) can thus be defined relative to the "excitement" and "pleasure" emotions by coordinates within the multi-dimensional continuous space. Thus, user interaction with the control can be used to define positions and corresponding coordinates that can act as weights for search criteria in order to perform a search of digital content using a single user input.

[0023] Continuing again with the emotion example above, a user input can be received via the control that specifies a position within the multi-dimensional continuous space defined using positive and negative amounts of excitement and pleasure. The user input can, for example, use the control together with the text input "dog" to specify a position corresponding to the emotion "relaxed". The position (e.g., the coordinates of the position) and the text input form a search query that is then used to locate digital content (e.g., digital images) that includes similar objects (e.g., by using tags) and is also associated with similar coordinates within the multi-dimensional continuous space.

[0024] For example, the position corresponding to "relaxed" specifies a medium positive amount of pleasure and a medium negative amount of excitement. In this way, this position is used to specify weights within a multi-dimensional continuous space defined by excitement and pleasure to define an emotion that would otherwise be difficult (if not impossible) to define using conventional techniques. Additionally, this overcomes the challenges of conventional tag-based methods that are based on determining the consistency between the intent of a user input when searching for digital content and the intent expressed by tags associated with the digital content.

[0025] Although digital images and emotions are described in this example, the control can be used to define various other search criteria as part of a multi-dimensional continuous space, such as digital content characteristics, such as creation settings (e.g., exposure, contrast), audio characteristics (e.g., timbre, range), etc. Additionally, these search techniques can be utilized to search various types of digital content, such as digital images, digital movies, digital audio, web pages, digital media, and so on. Further discussion of these and other examples is included in the following sections and illustrated using the corresponding figures.

[0026] In the following discussion, an example environment in which the search techniques described herein can be employed is first described. Example processes that can be executed in the example environment as well as other environments are also described. Thus, the execution of the example processes is not limited to the example environment, and the example environment is not limited to the execution of the example processes.

[0027] Example environment

[0028] Figure 1 FIG. 14 is a diagram of a digital media search environment 100 in an example implementation that is operable to employ the digital content search techniques described herein. The illustrated environment 100 includes a computing device 102 communicatively coupled to a service provider system 104 via a network 106 such as, for example, the Internet. The computing devices implementing the computing device 102 and the service provider system 104 can be configured in a variety of ways.

[0029] For example, the computing device can be configured as a desktop computer, a laptop computer, a mobile device (e.g., assuming a handheld configuration such as a tablet computer or a mobile phone as shown), etc. Thus, the range of computing devices can range from full-resource devices with substantial memory and processor resources (e.g., personal computers, gaming consoles) to low-resource devices with limited memory and / or processing resources (e.g., mobile devices). Additionally, the computing device can represent multiple different devices, such as multiple servers used by an enterprise to perform the operations illustrated for the service provider system 104 and further described with respect to Figure 11 Although the search techniques are shown and described as occurring over the network 106 in this example, these techniques can also be implemented locally by the computing device 102 alone.

[0030] Computing device 102 is shown as including a communication module 108 configured to communicate with a service provider system 104 via network 106. Communication module 108 may be configured as a browser, a network-enabled application, a plug-in module, etc. Communication module 108 includes a search input / output (I / O) module 110 configured to generate a search query 112 for searching digital content and output search results 114 resulting from the search in user interface 116.

[0031] User interface 116 in the illustrated example includes a text input portion 118 through which user input for specifying text as part of search query 112 (e.g., "dog") can be received. User interface 116 also includes a control 120 that includes a representation of a multi-dimensional continuous space that in this example is defined relative to a first criterion 122 associated with a first axis of control 120 and a second criterion 124 associated with a second axis of control 124, for example perpendicular to each other. Both first criterion 122 and second criterion 124 can be defined using positive, neutral, and negative amounts, as further described below. The space is continuous because corresponding positions within the space together define a respective amount for each search criterion among the search criteria. Thus, a single user input 126 can be used to define a position relative to both the first and second axes together and to define the corresponding weights of these axes.

[0032] Search query 112 including text and position is shown as being transmitted from computing device 102 to a digital content search system 128 of service provider system 104 via network 106. Digital content search system 128 is configured to search digital content 130 based on search query 112 and generate search results 114 therefrom for transmission back to computing device 102. Although digital content 130 is shown as being locally stored by storage device 132 of service provider system 104, digital content 130 can be maintained elsewhere by a third party system, for example.

[0033] The digital content search system 128 includes a multi-dimensional search module 134 that represents a function of supporting the search of digital content 130 by utilizing a multi-dimensional continuous space represented by the control 120. For example, each digital content item can be associated with a position (e.g., coordinates) within the multi-dimensional digital space. Thus, the multi-dimensional search module 134 can incorporate the relationship between the position specified by the search query 112 with respect to the space and the position specified for the corresponding digital content 130 item. In this way, the digital content search system 128 can support increased granularity and flexibility as part of searching for digital content 130, especially for concepts that are difficult to express in words, such as emotions.

[0034] Figure 2 An example of the control 120 configured to support a multi-dimensional continuous space for searching using emotions is depicted. Figure 1 The control 120 supports user input for continuously specifying the intensity of at least two search criteria, which in this case are the emotional signals of pleasure (P) and excitement (E). This is achieved by mapping the "P" and "E" parameters to the "X" and "Y" axes on a multi-dimensional continuous space (in this example, a two-dimensional (2D) plane). To specify a combination of "P" and "E", user input for specifying a position with respect to this representation of the 2D plane is received, for example, via gestures (e.g., tagging, dragging) of the cursor control device shown, spoken words received via the user interface, etc. For example, user input for specifying a position as a pin can be received, where the coordinates of the position are displayed in the user interface.

[0035] To further enhance the user experience and facilitate user intuition about the meaning of the position (i.e., coordinates), text labels are displayed as part of the control 120 that indicate fine-grained emotions corresponding to the respective parts of the 2D plane. The example shown includes excitement, elation, happiness, contentment, relaxation, calmness, tiredness, boredom, frustration, disappointment, anger, and tension. Each of these fine-grained emotions corresponds to a respective amount of "P" and "E", which can be positive, neutral, or negative. For example, excitement, elation, and happiness are marked in the upper right region of the 2D plane, which map to instances where both the "P" and "E" signals are positive. Similarly, frustration, boredom, and tiredness are marked in the lower left region to indicate a relatively negative amount of both the "P" and "E" signals. In this way, user input can be effectively provided to support digital search, the further discussion of which is included in the following sections and shown in the corresponding figures.

[0036] Generally, the functions, features, and concepts described in the examples above and below can be adopted in the context of the example processes described in this section. Additionally, the functions, features, and concepts described with respect to the different figures and examples in this document can be interchanged with one another and are not limited to implementation within the context of a specific figure or process. Moreover, the blocks associated with the different representative processes and corresponding figures in this document can be applied and / or combined in different ways. Accordingly, the various functions, features, and concepts described with respect to the different example environments, devices, components, figures, and processes in this document can be used in any suitable combination and are not limited to the specific combinations represented by the examples enumerated in this specification.

[0037] Multi-dimensional digital content search

[0038] Figure 3 illustrates system 300 in an example implementation, which more particularly shows the operation of search I / O module 110 and digital content search system 128 when performing a multi-dimensional digital content search Figure 1 thereof. Figure 4 illustrates example 400 of a multi-dimensional digital content search involving sentiment. Figure 5 illustrates another example 500 of a multi-dimensional digital content search involving sentiment. Figure 6 illustrates process 600 in an example implementation, where a control including a representation of a multi-dimensional continuous space is utilized as part of a digital content search.

[0039] The following discussion describes search techniques that can be implemented using the previously described systems and devices. Aspects of the process can be implemented in hardware, firmware, software, or a combination thereof. The process is shown as a set of blocks specifying operations to be performed by one or more devices, and is not necessarily limited to the order shown for operations to be performed by the corresponding blocks. In the following discussion sections, reference may be made interchangeably to Figures 1 to 6 .

[0040] First, in this example, as Figure 3 shown, search I / O module 110 includes user interface module 302 and search query generation module 304. User interface module 302 is configured to output Figure 1 user interface 116. As part of it, user interface module 302 includes text input module 306 configured to receive user input for specifying text 308, for example, via text input section 118. User interface module 302 also includes control module 310 configured to display control 120 in user interface 116.

[0041] Control 120 includes a representation of a multi-dimensional continuous space, including as Figure 1A first axis associated with a representation of a first search criterion and a second axis associated with a representation of a second search criterion are shown (block 602). As Figure 2 shown, the first search criterion and the second search criterion can respectively correspond to emotions, such as pleasure and excitement.

[0042] Then, user input is received through interaction with control 120. The user input provides an indication 312 of a position 314 (e.g., coordinates 316) defined relative to a multi-dimensional continuous space. The user input also includes text 308 (block 604). For example, text 308 can be received through a text input section 118 output by a text input module 306, such as the word "dog" entered using a keyboard, a spoken utterance, a gesture, etc. An indication 312 can also be received that specifies a position 314 (e.g., coordinates 316) defined relative to a representation of the multi-dimensional continuous space defined by control 120, e.g., "clicking" on a position by using a cursor control device, a click gesture, etc.

[0043] As Figure 4 shown in example 400, for example, the search query 112 can include the text 308 "girl". The search query 112 also includes coordinates 322 defined relative to a representation of the multi-dimensional continuous space of control 120 output by control module 310, in this case, the coordinates 322 indicating a position near "excitement" and "delight" to indicate a high amount of "excitement" and a medium amount of "pleasure". On the other hand, in Figure 5 example 500, the search query 112 includes the text 308 "boy". The search query 112 also includes coordinates 322 defined relative to the multi-dimensional continuous space of control 120 output by control module 310, the coordinates 322 indicating a position near "boredom" and "tiredness" to indicate a relatively low amount of "excitement" and a negative amount of "pleasure". Thus, in both cases, the coordinates 322 specify the weights to be applied to the two emotions by a single user input, and the weights can be positive or negative.

[0044] Then, the user interface module 302 outputs the text 308 and the indication 312 to the search query generation module 304. The search query 112 is generated by the search query generation module 304 based on the position 314 from the user input (e.g., coordinates 316 relative to the multi-dimensional continuous space) and the text 308 (block 606). Then, the search query 112 is transmitted to and received by the search query collection module 318 of the digital content search system 128 (block 608). As previously described, this can be performed remotely using network 106 or locally at a single computing device 102.

[0045] The multi-dimensional search module 134 generates search results 114 using the search queries 112 collected by the search query collection module 318. The search results 114 are based on the search of multiple digital contents 130 by the machine learning model 320 based on the text 308 and location 314 from the search query 112 (block 610). The machine learning model 320 can, for example, be configured as an integrated model, as further described with respect to Figure 7 which includes an image model and a label-based model. The integrated model can thus be used to generate coordinates 322 for corresponding items of the digital content 130. In this way, the coordinates 316 of the indication 312 of the text 308 and location 314 from the search query 112 can be used to locate digital contents 130 with similar text and coordinates. Then, the search results 114 are output by the output module 324 (block 612). In this way, the multi-dimensional search module 134 supports increased flexibility and granularity compared to conventional techniques.

[0046] Continuing Figure 4 with the first example 400, the search query 112 can include the text 308 "girl". The search query 112 also includes coordinates 322 defined relative to the representation of the multi-dimensional continuous space of the controls 120 output by the control module 310, the coordinates 322 indicating a position near "excited" and "joyful" to define a relatively high positive amount of "excitement" and a medium positive amount of "pleasure". Thus, the multi-dimensional search module 134 generates search results 114 which, in this example, show a girl with a high amount of excitement and a medium amount of pleasure based on the coordinates assigned to the digital image, for example, a girl raising her hand and jumping off a dock.

[0047] Similarly, in Figure 5 example 500, the search query 112 includes the text 308 "boy". The search query 112 also includes coordinates 322 defined relative to the multi-dimensional continuous space of the controls 120 output by the control module 310, the coordinates 322 indicating a position near "bored" and "tired". This indicates a relatively low negative amount of "excitement" and a low negative amount of "pleasure". Thus, the multi-dimensional search module 134 generates search results 114 which include a digital image associated with the text 308 "boy" and coordinates 322 which show a boy exhibiting a low amount of excitement and pleasure, for example, a boy lying on a couch staring at a tablet computer. As a result, the multi-dimensional continuous space supports search techniques with potentially higher computational efficiency and accuracy. Further discussion of implementation examples is included in the following sections and is illustrated using corresponding figures, which include additional details related to the configuration of the digital content to support multi-dimensional continuous search and the use of digital content as part of the search.

[0048] Implementation example

[0049] In this implementation example, the control 120 is configured to support emotion-based digital image search. Emotion-based image search is a powerful tool that can be used by a computing device to find digital images that trigger corresponding emotions. For example, different digital images may evoke different emotions in humans. In this case, the "pleasure" and "excitement" emotions are used as a basis for defining other emotions by using a multi-dimensional continuous space.

[0050] Conventional search solutions are based on a tag-based approach, where the search is limited to a single emotion that is part of the search query, such as "happy child" or "angry child". For example, if there is a single emotion for a subject, the conventional tag-based search works well, but it is not very good in terms of granularity and flexibility. For example, "happy child", "sad girl" work well in tag-based search, however, conventional techniques do not support searching for multiple items together (such as "happy enthusiastic child" or "sad calm girl") with acceptable accuracy. Other conventional techniques do not support the ability to attach weights to items expressing emotions, nor can they be performed together. For example, conventional techniques do not support specifying the weights of happiness or enthusiasm in a search such as "happy enthusiastic child".

[0051] Therefore, the techniques described herein support the ability to search for digital images having different degrees of emotion associated with the digital images. Accordingly, these techniques support the user experience with improved efficiency and accuracy for performing a search of digital images, as further described below. As previously mentioned, the multi-dimensional search module 134 supports the search by leveraging a multi-dimensional continuous space. In this example, this space is used to conceptualize and define human emotions by defining the positions of these emotions within the space (e.g., in a two-dimensional grid).

[0052] Figure 7 is more detailedly depicted Figure 3 An example implementation 700 of the machine learning model 320 of the multi-dimensional search module 134. In this example, the machine learning model 320 is implemented as an ensemble model 702, and the ensemble model 702 includes an image-based model 704 and a tag-based model 706.

[0053] The image-based model 704 is trained in two stages. First, the base model is trained using training data 708 from the base dataset 710 based on a relatively large number of weakly supervised digital images. Then, the base model is "fine-tuned" using the fine-tuned dataset 712 to generate the image-based model 704.

[0054] In this example, the base model of the image-based model 704 is formed using the Resnet50 architecture. Training a machine learning model to recognize emotions in digital images involves a large dataset. To address this, a large-scale base dataset 710 with weak derivations is curated, which includes over a million digital images covering various emotion concepts related to humans, scenes, and symbols. A part of the base dataset 710 may be incomplete and noisy, for example, the digital images include few labels or incomplete labels or labels that are not relevant or loosely related to the digital images. Since the representations of visual data and text data need to be semantically close to each other, the labels and the relevant information in the digital images serve to regularize the image representation. Thus, in this example, training is performed on the joint text and visual information of the digital images.

[0055] The base dataset 710 uses 690 emotion-related labels as labels to give a diverse set of emotion labels, thus avoiding the difficulty of manually obtaining emotion annotations. The base dataset 710 is used to train the feature extraction network of the image-based model 704, which is further regularized using joint text and visual embeddings and text distillation. The model provides 690-dimensional probability scores for 690 labels (main task) and 300-dimensional feature vectors (main task). 8-dimensional probability scores for 8 categories (auxiliary task) are also trained. For the above three tasks, a multi-task loss is used to train the model.

[0056] For the fine-tuned dataset 721, 21,000 digital images are collected, and each digital image is labeled for 25 values in -2, -1, 0, +1, +2 based on two search criteria (e.g., two axes) in each dimension. This annotation is performed independently along each axis. To fine-tune the base model using the fine-tuned dataset 712, the last layer of the base model is removed, and a fully connected layer is added to the head of the base model, where the output is mapped to a class with two scores. The multi-class log loss is used to train the model as follows:

[0057]

[0058] For the label-based model 706, the inventory dataset of the training data 708 includes 140 million digital images with weak labels (e.g., text labels provided at least partially by users). Each digital image also includes a variable number of labels. To find the coordinates for each digital image within a multi-dimensional continuous space, coordinates are assigned to each emotion label in the emotion labels based on this space, for example, using 2D axes based on its position on a 2D grid.

[0059] In Figure 8In the illustrated example 800, for example, emotional label coordinates can be defined with respect to the X-axis of pleasure and the Y-axis of excitement. For example, emotions and corresponding coordinates can include the following:

[0060] · Happy [0.67, 1]

[0061] · Ecstatic [0.67, 0.67]

[0062] · Excited [.33, 1]

[0063] · Nervous [-0.33, 1]

[0064] · Angry [-0.67, 0.67]

[0065] · Disappointed [-1, 0.33]

[0066] · Depressed [-1, -0.33]

[0067] · Bored [-0.67, -0.67]

[0068] · Tired [-0.33, -1]

[0069] · Calm [0.33, -1]

[0070] · Relaxed [0.67, -0.67]

[0071] · Satisfied [0.67, -0.33]

[0072] Therefore, consider Figure 9 Example 900, where the digital image 902 includes the following labels 904.

[0073] · Happy

[0074] · Children

[0075] · Parents

[0076] · Sunshine

[0077] · Joy

[0078] · Meadow

[0079] · Relaxed

[0080] · Playing

[0081] · Evening

[0082] · Sky

[0083] · Trees

[0084] · Cover

[0085] · Daylight

[0086] ·Mother

[0087] ·outdoor

[0088] In this example, digital image 902 is associated with 15 tags. However, among these tags, three tags (1) happy, (2) joyful, and (3) relaxed represent emotions. Therefore, coordinates may be assigned to digital image 902 for each of these tags and / or as a whole.

[0089] For example, for the entire digital image 902, first, the tag associated with the digital image 902 is combined with the tag from Figure 8 The labels of the examples are matched (e.g., using natural language processing, vectors in word2vec space, etc.) and the corresponding coordinates are obtained. For example, the emotions "happy" and "joy" can be mapped to the "happy" label in the table. Similarly, the emotion "relaxed" can be mapped to "relaxed" in the table.

[0090] Next, the coordinates "[[0.67, 1]" corresponding to "happy" and the coordinates "[0.67, -0.67]" corresponding to "relaxed" are obtained. Then, the coordinates of the entire digital image 902 are calculated as the average value of the coordinates [(0.67+0.67) / 2, (1+(-0.67)) / 2] = [0.67, 0.16]. The resulting coordinates [0.67, 0.16] are assigned as the position of the digital image 902 in the multi-dimensional continuous space. Therefore, in this case, the digital image 902 is located somewhere in the first quadrant.

[0091] Likewise, consider Figure 10 1000 , where a digital image 1002 includes the following tag 1004 .

[0092] ·boring

[0093] ·happy

[0094] Calm

[0095] ·family

[0096] ·spouse

[0097] Here, three out of five labels are associated with emotions, namely, "bored," "happy," and "calm." These emotions correspond to coordinates [-0.67, -0.67], [0.67, 0.67], and [0.33, -1], respectively. Therefore, the coordinates associated with the digital image 1002 as a whole can be calculated as follows: [((-.067)+(0.67)+(0.33)) / 3, ((-0.67)+(0.67)+(-1)) / 3]=[0.11, -0.33] Therefore, in this case, the digital image 10002 is located somewhere in the fourth quadrant.

[0098] The image-based model 704 and the label-based model 706 form an integrated model 702 adopted by the multi-dimensional search module 134. In one example, equal weights are assigned to the two models, and thus the final model is represented as M.

[0099] M = 1*m1+(1 - 1)*m2

[0100] where "m1" is the image-based model 704, "m2" is the label-based model 706, and l = 0.5, which has been found in practice to provide the best results.

[0101] The output of the Resnet-based image model is [0.75, 0.67], and the output of the label-based model is [0.67, 0.16]. The output of the integrated model 702 for l = 0.5 can be calculated as 0.5*[0.75, 0.67]+(1 - 0.5)*[0.67, 0.16]=[0.71, 0.41]. Some digital images in the training dataset may not include sentiment labels. In such cases, l = 1 is assigned, and the output of the integrated model becomes

[0102] M = m1

[0103] where "m1" is the Resnet-based image model. The output of the integrated model 702 is a score in the format of [x, y], where the scores on both the X and Y axes are between [-1, 1]. These [x, y] coordinates also correspond to points in a multi-dimensional continuous space.

[0104] The multi-dimensional search module 134 can adopt an elastic search index, where coordinates are generated offline to support real-time operations when receiving a search query 112 to generate a search result 114. For this purpose, the infrastructure of the multi-dimensional search module 134 can include an analyzer and an elastic search index. The analyzer is used as part of the setup, where the integrated model is deployed as a web service inside a Docker container. Additionally, the analyzer can be scaled to allocate sufficient resources to index millions of digital images in a short time.

[0105] The elastic search index is an elasticsearch-based index that can be queried to return digital content 130 (e.g., digital images) that is closest to the location specified as part of the search query 112 based on the L2 distance. To create the index, a product quantization technique is used, which involves compressing the feature embeddings, performing bucketing (clustering), and assigning to one of 1k buckets. The pre-established reverse ES index allows for real-time retrieval of digital content 130.

[0106] To compress the size of the feature vector of an image and compute the PQ code, the following operations can be performed. First, the embedding space is subdivided into subspaces of each 8 bits. Each byte represents the bucket identifier of the elastic search index. From the perspective of the search for the nearest neighbor, each byte represents the centroid of the clusters in the KNN. Then, each embedded subspace vector is encoded using the ID of the nearest cluster (bucket). The PQ code is computed using the subspace ID, and the PQ code and the bucket ID are stored in the elastic search as an inverted index.

[0107] Once the inverted ES index is established. The results can be retrieved through the following mechanism.

[0108] 1. The user makes a query using a 2D grid;

[0109] 2. The mentioned analyzer transforms the query and the output is sent to the PQ-codes plugin;

[0110] 3. The PQ code plugin compares the input vector with the subspace ID and returns the closest subspace ID based on the L2 distance. This is an example of "approximate nearest neighbor" search.

[0111] 4. The digital content 130 from the (multiple) buckets associated with the subspace ID is used to generate the search result 114; and

[0112] 5. The inverted index can be used to limit the search to the N closest buckets.

[0113] In this way, real-time search can be implemented as part of the multi-dimensional digital content search technology described herein.

[0114] For example, in a situation where 180 million digital images can be processed (e.g., as part of an inventory digital image service), certain regions of the multi-dimensional continuous space may be dense while other regions may be sparse. Therefore, to improve the operational efficiency of the computing device performing the search, this can be achieved by not directly searching for the closest digital images in the space. For example, searching for "happy children" can yield ten million digital images as part of the search result 114. Therefore, to improve the processing efficiency, the positions of the digital images within the multi-dimensional continuous space are pre-computed and clustered into bins and the search is performed based on these bins (e.g., centroids).

[0115] The multi-dimensional continuous space (e.g., Figure 2The 2D space shown (e.g., 1000) can be divided into multiple boxes, and the top “X” digital images within the box are positioned to improve efficiency as part of local neighbor search. Additionally, the search results 114 output in the user interface 116 can include a density map to show “thing areas” relative to a representation of a multi-dimensional continuous space, e.g., as an availability heat map. Further, based on the amount of digital content assigned to the area, the grid size can vary in areas for representing different emotions and can support “zooming” to support different levels of granularity. Other examples can also be envisioned without departing from the spirit and scope of the present invention.

[0116] Example systems and devices

[0117] Figure 11 An example system generally designated 1100 includes an example computing device 1102. The example computing device 602 represents one or more computing systems and / or devices that can implement the various techniques described herein. This is illustrated by including a multi-dimensional search module 134. The computing device 1102 can be, for example, a server of a service provider, a device associated with a client (e.g., a client device), a system-on-chip, and / or any other suitable computing device or computing system.

[0118] The example computing device 1102 shown in the figure includes a processing system 1104, one or more computer-readable media 1106, and one or more I / O interfaces 1108 communicatively coupled to each other. Although not shown, the computing device 1102 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus can include any one or combination of different bus structures using any of the various bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus. Various other examples are also envisioned, such as control and data lines.

[0119] The processing system 1104 represents the functionality to perform one or more operations using hardware. Thus, the processing system 1104 is shown as including hardware elements 1110 that can be configured as a processor, functional blocks, etc. This can include being implemented in hardware as an application-specific integrated circuit or other logic device formed using one or more semiconductors. The hardware elements 1110 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can include (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, processor-executable instructions can be electronically executable instructions.

[0120] The computer-readable size medium 1106 is shown as including a memory / storage device 1112. The memory / storage device 1112 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage component 1112 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic discs, etc.). The memory / storage component 1112 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). The computer-readable medium 1106 may be configured in various other ways as further described below.

[0121] (Multiple) input / output interfaces 1108 represent the functionality that allows a user to input commands and information into the computing device 1102 and also allows information to be presented to the user and / or other components or devices using various input / output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone, a scanner, a touch function (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., which may use visible or non-visible wavelengths such as infrared frequencies to recognize movement not involving touch as a gesture), etc. Examples of output devices include a display device (e.g., a monitor or a projector), a speaker, a printer, a network card, a haptic response device, etc. Thus, the computing device 1102 may be configured in various ways as further described below to support user interaction.

[0122] Various techniques may be described herein in the general context of software, hardware elements, or program modules. Generally, such modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The terms "module", "function", and "component" as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, which means that these techniques may be implemented on various commercial computing platforms having various processors.

[0123] The implementation of the described modules and techniques may be stored on or transmitted through some form of computer-readable medium. The computer-readable medium may include various media that may be accessed by the computing device 1102. By way of example and not limitation, the computer-readable medium may include "computer-readable storage media" and "computer-readable signal media".

[0124] "Computer-readable storage medium" can refer to a medium and / or device that can persistently and / or non-transiently store information, as opposed to merely signal transmission, carrier waves, or signals themselves. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data, such as volatile and non-volatile, removable and non-removable media and / or storage devices. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical memories, hard disks, magnetic tape cartridges, tapes, magnetic disk memories, or other magnetic storage devices, or other storage devices, tangible media, or articles suitable for storing the desired information and accessible by a computer.

[0125] "Computer-readable signal medium" can refer to a signal-bearing medium configured to transmit instructions to the hardware of computing device 1102, such as via a network. A signal medium typically can contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. A signal medium also includes any information delivery medium. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0126] As previously described, hardware elements 1110 and computer-readable medium 1106 represent modules, programmable device logic, and / or fixed device logic implemented in hardware that can be used to implement at least some aspects of the techniques described herein, such as executing one or more instructions. Hardware can include integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implemented components of silicon or other hardware. In this context, the hardware can operate as a processing device that executes program tasks defined by instructions and / or logic implemented in hardware, as well as hardware for storing instructions for execution (e.g., the previously described computer-readable storage medium).

[0127] The foregoing combinations can also be used to implement the various techniques described herein. Thus, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic included on a computer-readable storage medium of a certain form and / or implemented by one or more hardware elements 1110. The computing device 1102 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, the implementation of the modules executable as software by the computing device 1102 can be at least partially implemented in hardware, for example, by using the computer-readable storage medium and / or the hardware elements 1110 of the processing system 1104. The instructions and / or functions can be executed / operated by one or more articles of manufacture (e.g., one or more computing devices 1102 and / or the processing system 1104) to implement the techniques, modules, and examples described herein.

[0128] The techniques described herein can be supported by various configurations of the computing device 1102 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part by using a distributed system, such as being implemented on the "cloud" 1114 via the platform 1116 as described below.

[0129] The cloud 1114 includes and / or represents a platform 1116 for resources 1118. The platform 1116 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud 1114. The resources 1118 can include applications and / or data that can be used when performing computer processing on servers remote from the computing device 1102. The resources 1118 can also include services provided via the Internet and / or via a subscriber network (such as a cellular or Wi-Fi network).

[0130] The platform 1116 can abstract resources and functionality to connect the computing device 1102 with other computing devices. The platform 1116 can also be used to abstract the scaling of resources to provide a corresponding scale level to meet the demand for resources 1118 implemented via the platform 1116. Thus, in an interconnected device embodiment, the implementation of the functionality described herein can be distributed throughout the system 1100. For example, the functionality can be implemented partially on the computing device 1102 and via the platform 1116 that abstracts the functionality of the cloud 1114.

[0131] Conclusion

[0132] Although the invention has been described in language specific to structural features and / or method acts, it is to be understood that the invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed invention.

Claims

1. A search method implemented by a computing device, the method comprising: Receiving, by the computing device, a search query via an input provided by a user, the search query comprising: A text query; and An indication of a position relative to a multi-dimensional continuous space defined using a first axis corresponding to a first sentiment and a second axis corresponding to a second sentiment; Searching, by the computing device, a plurality of digital images based on the text query and the indication of the position, the search being performed using a machine learning model configured as an ensemble model, the ensemble model comprising an image-based model and a label-based model, the image-based model being trained on joint text and visual information, and the label-based model being trained on digital images each having one or more sentiment-based labels; Generating, by the computing device, search results based on the search; and Outputting, by the computing device, the search results.

2. The method according to claim 1, wherein the first axis corresponds to excitement or enthusiasm, and the second axis corresponds to pleasure or happiness.

3. The method according to claim 1, wherein the first axis and the second axis respectively define positive and negative amounts for the first sentiment and the second sentiment within the multi-dimensional continuous space.

4. The method according to claim 1, wherein the indication of the position specifies weights respectively assigned to the first sentiment and the second sentiment.

5. The method according to claim 1, wherein each of the one or more sentiment-based labels has assigned coordinates within the multi-dimensional continuous space.

6. The method according to claim 1, wherein the indication is generated by receiving user input via a control output in a user interface, the user input indicating the position relative to a representation of the multi-dimensional continuous space displayed as part of the control output.

7. The method according to claim 1, wherein the indication of the position is specified using coordinates relative to the multi-dimensional continuous space.

8. The method according to claim 7, wherein the multi-dimensional continuous space comprises at least two dimensions.

9. A search system, the system comprising: A search query collection module, at least partially implemented in hardware of a computing device to receive a search query via an input provided by a user, the search query comprising: A text query; and Coordinates specified relative to a multi-dimensional continuous space; A multi-dimensional search module, at least partially implemented in hardware of the computing device to generate search results based on a search of a plurality of digital images, the plurality of digital images based on the search query using a machine learning model configured as an ensemble model, the ensemble model comprising an image-based model and a label-based model, the image-based model being trained on joint text and visual information, and the label-based model being trained on digital images each associated with one or more sentiment-based labels, the one or more sentiment-based labels having assigned coordinates within the multi-dimensional continuous space; and An output module, at least partially implemented in the hardware of the computing device, to output the search results.

10. The system according to claim 9, wherein during training, the ensemble model assigns positions in the multi-dimensional continuous space for each of the digital images based on the assigned coordinates of the one or more sentiment-based tags for each of the digital images.

11. The system according to claim 9, wherein the multi-dimensional continuous space defines corresponding amounts of at least two sentiments.

12. The system according to claim 9, wherein a first axis and a second axis respectively define positive and negative amounts for a first search criterion and a second search criterion within the multi-dimensional continuous space.

13. The system according to claim 9, wherein the coordinates specify weights respectively assigned to a first sentiment and a second sentiment within the multi-dimensional continuous space.

14. The system according to claim 9, wherein the coordinates are generated by receiving user input via a control output in a user interface, the user input indicating the position of the coordinates relative to a representation of the multi-dimensional continuous space displayed as part of the control output.

15. A method implemented by a computing device for generating a plurality of coordinates, the method comprising: receiving, by the computing device, a plurality of digital images; generating, by the computing device, for each of the plurality of digital images, a plurality of coordinates relative to a multi-dimensional continuous space, the generating including: respectively locating a plurality of sentiment-based tags associated with the plurality of digital images; generating, for each of the plurality of sentiment-based tags, a plurality of coordinates; and calculating, for the plurality of digital images as a whole, the plurality of coordinates based on the generated plurality of coordinates of the sentiment-based tags for each of the digital images; and training, by the computing device, a machine learning model based on: the plurality of digital images; and the plurality of coordinates of the plurality of digital images as a whole.

16. The method according to claim 15, wherein the machine learning model is an ensemble model.

17. The method according to claim 16, wherein the ensemble model includes an image-based model and a tag-based model.

18. The method according to claim 15, further comprising: using the trained machine learning model to search the plurality of digital images based on a search query; and outputting the search results of the search.

19. The method according to claim 18, wherein the search query includes a digital image and coordinates.

Citation Information

Patent Citations

  • Providing a response in a session

    CN110301117A

  • Electronic device and method for controlling the electronic device

    CN110998565A

  • Systems and methods for mixed-media content guidance

    US20120272185A1