Information processing device, information processing method, and program
By allowing users to set weights for query words, the information processing device enhances search relevance by generating weighted feature vectors, addressing the challenge of inaccurate search results in existing data search technologies.
Patent Information
- Application Number
- PCT/JP2025/012367
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-16
AI Technical Summary
Existing data search technologies often produce results that do not accurately reflect users' intentions, especially for users unfamiliar with data search processes, due to the difficulty in generating optimal keywords or queries.
An information processing device and method that allows users to set weights for each word or word string in a query, generating a weighted query-related feature vector to enhance similarity analysis and improve the accuracy of search results.
The solution enables search results that more accurately reflect users' intentions by analyzing the similarity between weighted query-related feature vectors and content-related feature vectors, thereby improving the relevance of search outcomes.
Smart Images

Figure JP2025012367_16102025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that make it easier to obtain search results that reflect a user's intentions in data search processing.
[0002] In data search processing targeting databases or data on the Internet, a user can perform a search by inputting a query (search statement) that is a keyword or a free-form statement expressed in text.
[0003] For example, in image search processing, a user generates a query (search statement) or keywords that describe the characteristics of the image they want to search for. For example, by entering "dog photos" as a query (search statement) and performing a data search, it is possible to select and retrieve image data of dog photos from a large number of image data available in databases or on the Internet.
[0004] Incidentally, for example, Patent Document 1 (Japanese Patent Laid-Open Publication No. 2008-243024) is a conventional technique that discloses data search processing using keywords.
[0005] However, data searches using keywords or queries (search statements) often produce search results that differ from the user's intention. In particular, users who are unfamiliar with data search processes have difficulty generating optimal keywords or queries (search statements), making it difficult to obtain the intended search results.
[0006] Japanese Patent Application Laid-Open No. 2008-243024
[0007] The present disclosure has been made in consideration of, for example, the above-mentioned problems, and aims to provide an information processing device, an information processing method, and a program that make it easier to obtain search results that reflect the user's intentions in data search processing.
[0008] A first aspect of the present disclosure is an information processing device that acquires a query composed of text to be applied to data search processing, and has a data processing unit that executes data search processing based on the query, wherein the data processing unit acquires, along with the query, weight information for each word or word string that constitutes the query, and executes search processing for data in a category different from text using the query and the weight information for each word or word string.
[0009] Furthermore, a second aspect of the present disclosure resides in an information processing method that acquires a query composed of text to be applied to data search processing and weight information for each word or word string that constitutes the query, and performs search processing for data in a category different from the text by utilizing the query and the weight information for each word or word string.
[0010] Furthermore, a third aspect of the present disclosure is a program for causing an information processing device to execute information processing, the program causing a data processing unit of the information processing device to acquire a query composed of text to be applied to data search processing and weight information for each word or word string that constitutes the query, and execute search processing for data in a category different from text using the query and the weight information for each word or word string.
[0011] The program of the present disclosure is, for example, a program that can be provided in a computer-readable format via a storage medium or a communication medium to an information processing device or a computer system capable of executing various program codes. By providing such a program in a computer-readable format, processing according to the program is realized on the information processing device or the computer system.
[0012] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description of the embodiments of the present disclosure and the accompanying drawings. Note that in this specification, a system refers to a logical collective configuration of multiple devices, and is not limited to devices that are located within the same housing.
[0013] According to an embodiment of the present disclosure, a device and method are realized that perform search processing for data in a category different from text using a query and word-by-word weight information. Specifically, for example, the device includes a data processing unit that inputs a query consisting of text to be applied to the data search processing and performs data search processing based on the query. The data processing unit inputs weight information for each word that constitutes the query along with the query, and performs search processing for data in a category different from text using the query and word-by-word weight information. For words assigned with a positive weight, the device and method perform search processing with an increased influence, and for words assigned with a negative weight, the device and method perform search processing for data in a category different from text using a query and word-by-word weight information. Note that the effects described in this specification are merely examples and are not limiting, and additional effects may also be present.
[0014] 1 is a diagram illustrating an overview of data search processing using a query (search sentence) that is a free-form sentence; FIG. 2 is a diagram illustrating a configuration example and input example of a UI (user interface) that can input a weight setting query; FIG. 3 is a diagram illustrating a configuration example and input example of a UI (user interface) that can input a weight setting query; FIG. 4 is a diagram illustrating a configuration example and input example of a UI (user interface) that can input a weight setting query; FIG. 5 is a diagram illustrating a configuration example and input example of a UI (user interface) that can input a weight setting query; FIG. 6 is a diagram illustrating a configuration example and input example of a UI (user interface) that can input a weight setting query; FIG. 1 is a diagram illustrating an overview of a deep learning model that uses an attention mechanism. FIG. 2 is a diagram illustrating a specific example of processing performed by a weight-setting query-compatible feature vector generation unit. FIG. 3 is a diagram illustrating an example of data stored in a database. FIG. 4 is a diagram illustrating a configuration example in which a content encoder is provided inside a data processing unit of an information processing device. FIG. 5 is a diagram illustrating a flowchart illustrating a sequence of processing performed by an information processing device of the present disclosure. FIG. 6 is a diagram illustrating an example of a query input in processing to extract a specific scene intended by a user from movie content. FIG. 7 is a diagram illustrating processing by a data processing unit in processing to extract a specific scene intended by a user from movie content.FIG. 1 is a diagram illustrating a flowchart for explaining a processing sequence of a process for extracting a specific scene intended by a user from movie content. FIG. 2 is a diagram illustrating an example of a query input in a process for extracting a specific scene intended by a user from sports video content. FIG. 3 is a diagram illustrating the processing of a data processing unit in the processing for extracting a specific scene intended by a user from sports video content. FIG. 4 is a diagram illustrating stored data in a motion data database. FIG. 5 is a diagram illustrating an example of a query input in a process for extracting motion data intended by a user from motion data. FIG. 6 is a diagram illustrating the processing of a data processing unit in the processing for extracting motion data intended by a user from motion data. FIG. 7 is a diagram illustrating an example of a hardware configuration of an information processing device of the present disclosure.
[0015] The information processing device, information processing method, and program of the present disclosure will be described in detail below with reference to the drawings. The description will be made according to the following items: 1. Overview of data search processing using a query (search statement) that is free text 2. Example configuration of a UI (user interface) that allows input of a weight setting query, and example input of a weight setting query 2-(a) Example configuration of a query input area that allows weight setting by inputting numerical values, and example input of a query 2-(b) Example configuration of a query input area that allows weight setting by inputting brackets, and example input of a query 2-(c) Example configuration of a query input area that allows weight setting using a scroll bar, and example input of a query 2-(d) Example configuration of a query input area that allows weight setting using a color display (heat map), and example input of a query 3. 3. Specific examples of data search processing using a weight setting query 3-(1) Specific example of data search processing using a weight setting query by numerical input 3-(2) Specific example of data search processing using a weight setting query using a scroll bar 3-(3) Specific example of data search processing using a query in a language other than English 3-(4) Specific example of data search processing with music data as the search target 4. Details of the configuration and processing of the information processing device disclosed herein 5. Sequence of processing performed by the information processing device disclosed herein 6. Application examples of data search processing using the information processing device disclosed herein 6-(1) Example of processing for extracting a specific scene intended by the user from movie content 6-(2) Example of processing for extracting a specific scene intended by the user from sports video content 6-(2) Example of processing for extracting a specific scene intended by the user from sports video content 7. Example hardware configuration of the information processing device disclosed herein 8. Summary of the configuration of the present disclosure
[0016] [1. Overview of Data Search Processing Using a Query (Search Statement) that is a Free-Text Sentence] First, with reference to FIG. 1 and subsequent figures, an overview of data search processing using a query (search statement) that is a free-text sentence will be described.
[0017] 1 is a diagram showing an example of a UI (user interface) displayed on a display unit of an information processing device of the present disclosure, for example, a UI (user interface) used for data search processing targeting databases or data on the Internet. The UI shown in Fig. 1 is an example of a data search UI displayed on a display unit of an information processing device of the present disclosure, such as a PC, a smartphone, or a tablet terminal.
[0018] When performing a data search, the user first inputs a query (search statement) in the [Query] area at the top left of the UI shown in FIG. 1, that is, in the query (search statement) input area 11, in step S01.
[0019] 1 shows an example in which the query "query="people running on the beach" is input. That is, this is an example in which "people running on the beach" is input as a query.
[0020] After completing the query input in the query input area 11, in step S02, the user operates (touches) the search (search start) icon 12. This user operation starts a data search for the data to be searched, for example, a specific database or data on the Internet.
[0021] When the search process is completed, images are displayed as search results in a search result display area 13 at the bottom of the UI shown in FIG.
[0022] In the example shown in the figure, the query entered in the query input area 11 is "people running on the beach," i.e., an image corresponding to "people running on the beach" is displayed in the search result display area 13 as a search result.
[0023] However, the search results displayed in the search result display area 13 may differ from what the user intended.
[0024] That is, the query (search statement) entered in the query input area 11 shown in FIG. 1 is "people running on the beach." It is expected that the search intentions of users who perform data searches using a query consisting of the above-mentioned string of multiple words will differ depending on the user.
[0025] For example, assuming that there are multiple users who have performed searches using the above query, it is assumed that each of the multiple users has the following search intent: User a's search intent = To search for "images of many people running on the beach" User b's search intent = To search for "images of one person running on a quiet beach" User c's search intent = To search for "images of people running on a busy beach" User d's search intent = To search for "images of people running on a calm beach" In this way, the images that each search user wants to obtain are often different, and the search results displayed in search result display area 13 shown in Figure 1 may differ from the users' intent.
[0026] In such a case, the user needs to repeatedly create a different query and perform a search again in order to obtain the search results that the user intends.
[0027] The present disclosure provides a solution to such problems. Specifically, an information processing device according to the present disclosure enables a search in which weights are set for each word constituting the query (search statement) or for each word string consisting of multiple words, even when a search is performed based on a single query (search statement), thereby making it easier to obtain search results that reflect the user's intentions.
[0028] [2. Example of a configuration of a UI (user interface) that allows input of a weight setting query, and an example of input of a weight setting query] Next, an example of a configuration of a UI (user interface) that allows input of a weight setting query, and an example of input of a weight setting query will be described.
[0029] As described above, the information processing device of the present disclosure enables a search based on a query (search statement) that is a free-form text, in which weights are set for each word that constitutes the query (search statement) or for each word string consisting of multiple words.
[0030] The weights of the constituent words and word strings of the query (search statement) are used, for example, to generate a feature vector to be applied to the search process. That is, the data search process executed in the information processing device of the present disclosure is performed, for example, by applying a similarity analysis of feature vectors. Specifically, the similarity between the query-corresponding feature vector generated based on the query and the content-corresponding feature vector corresponding to each piece of content such as images to be searched is analyzed, and content (e.g., images) having content-corresponding feature vectors that are highly similar to the query-corresponding feature vector are output as search results.
[0031] Although the data search process based on the similarity analysis of feature vectors is an existing technology, the information processing device disclosed herein not only performs data search using a general "query-related feature vector" generated based on a query (search statement) composed of unweighted words, but also performs data search using a new feature vector that reflects the weights of the constituent words of the query (search statement) set by the user, i.e., a "weighted query-related feature vector."
[0032] The information processing device disclosed herein analyzes the similarity between the "weighted query-corresponding feature vector" and content-corresponding feature vectors corresponding to each piece of data to be searched, such as content such as images, and outputs content (images, etc.) having content-corresponding feature vectors that are highly similar to the "weighted query-corresponding feature vector" as search results.
[0033] This process makes it possible to obtain search results that reflect the user's intention with a higher degree of accuracy. A specific example of the process of generating the "weighted query-compatible feature vector" will be described in detail later.
[0034] 2 and subsequent figures will be used to describe a configuration example of a UI (user interface) that allows input of a weight setting query in an information processing device of the present disclosure, and an input example of a weight setting query. The information processing device of the present disclosure is, for example, an information processing device such as a PC, a smartphone, or a tablet terminal.
[0035] The display unit of the information processing device of the present disclosure displays a data search UI (user interface) similar to that described above with reference to Fig. 1. That is, a UI (user interface) used for data search processing targeting databases and data on the Internet is displayed.
[0036] The difference between the data search processing UI displayed on the display unit of the information processing device of the present disclosure and conventional UIs is the configuration of the [Query] area, i.e., the query (search statement) input area 11. The data processing unit of the information processing device of the present disclosure displays on the display unit the query input area 11 configured so that the user can set a "weight" for each constituent word (including word string) of the query (search statement). The "weight" is reflected in the generation of the above-mentioned "weight-set query-corresponding feature vector." Specifically, the larger the "weight" set for a word, the greater its influence on the data search.
[0037] A configuration example of the query input area 11 in the data search UI displayed on the display unit of the information processing device of the present disclosure and an input example of a weight setting query to the query input area will be described with reference to FIG. 2 and subsequent figures.
[0038] The following describes the configuration of the following four types of query input areas and examples of query input: (a) A configuration of a query input area that allows weight setting by inputting a numerical value, and an example of query input; (b) A configuration of a query input area that allows weight setting by inputting brackets, and an example of query input; (c) A configuration of a query input area that allows weight setting using a scroll bar, and an example of query input; (d) A configuration of a query input area that allows weight setting using a color display (heat map), and an example of query input.
[0039] (2-(a) Configuration of a query input area that allows weight setting by inputting numerical values and an example of a query input) First, a configuration of a query input area that allows weight setting by inputting numerical values and an example of a query input will be described.
[0040] FIG. 2 is a diagram illustrating the configuration of a [Query] area in a data search UI displayed on the display unit of the information processing device of the present disclosure, that is, a query input area, and an example of query input.
[0041] FIG. 2 shows three types of input examples of queries (search statements): (1) a normal search query; (a1) Example 1 of a search query with weight set by numerical input; and (a2) Example 2 of a search query with weight set by numerical input.
[0042] The example shown in "(1) Normal Search Query" is an example in which a conventional query, i.e., the following query, is entered in the query input area 11 in the data search UI, as described above with reference to FIG. 1 .
[0043] The example shown in "(a1) Example 1 of a weighted search query using numerical input" is an example of the configuration of the query input area 11a in the data search UI displayed on the display unit of the information processing device of the present disclosure and an example of a query entered in the query input area 11a.
[0044] In the query input area 11a, a query in which a "weight" is set for each constituent word of the query can be input. For example, the following weight-set query can be input: Query="(people:2.0) running on the beach"
[0045] In the above query, "(people:2.0)" means that the weight of one word that makes up the query, "people," is set to 2.0 times. The above query includes five words: "people," "running," "on," "the," and "beach."
[0046] Query="(people:2.0) running on the beach" This weight setting query sets the weight of "people" to 2.0 times, and the weights of "running," "on," "the," and "beach" to 1.0 times.
[0047] When this weight-setting query is input, the information processing device of the present disclosure executes a data search process in which the weight of "people" is set to twice that of the other four words, "running," "on," "the," and "beach." That is, the information processing device executes a data search process in which the influence of "people" is set to twice that of the other four words.
[0048] Specifically, the information processing device of the present disclosure generates a "weighted query-related feature vector" in which the influence of "people" is set to twice the influence of the other four words, and executes a similarity determination process between the generated "weighted query-related feature vector" with a high influence of "people" and a content-related feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, search results with a high influence of "people" are obtained.
[0049] "(a2) Example 2 of a search query with weighting set by numerical input" is another example of a query to be input into the query input area 11a in the data search UI displayed on the display unit of the information processing device of the present disclosure. That is, this shows an input example of the query: Query="(people: 2.0) running on the (beach: 3.0)."
[0050] "(people:2.0)" means that the weight of one word that makes up the query, "people," is set to 2.0 times. "(beach:3.0)" means that the weight of one word that makes up the query, "beach," is set to 3.0 times.
[0051] When this weight-set query is input, the information processing device of the present disclosure executes a data search process in which the weight of "people" is set to 2.0 times and the weight of "beach" is set to 3.0 times. That is, the data search process is executed in which the relative search influence of "people" with respect to the three unweighted words (running on the) is set to 2 times and 3 times, respectively.
[0052] Specifically, the information processing device of the present disclosure generates a "weighted query-related feature vector" in which the influence of "people" is set twice as strong and the influence of "beach" is set three times stronger than the three unweighted words (running on the), and then performs a similarity determination process between the generated "weighted query-related feature vector" and a content-related feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, search results with high influence of "people" and "beach" can be obtained.
[0053] The example of the weight setting query described with reference to Fig. 2 is an example in which a numerical value indicating a weight is associated with each word and input. The range of the weight value set for each word is a predetermined range, such as weight = -1.0 to +3.0. This weight numerical range may be a predetermined range, or may be freely set by the user.
[0054] A word that is set with a weight of 0 or less, such as -1.0, is a word that has less influence than other words in the search process. For example, query = "(people: -1.0) running on the beach" This weighted query sets the weight of "people" to -1.0, and the weights of "running," "on," "the," and "beach" to 1.0.
[0055] When this weight-set query is input, a data search process is executed in which the weight of "people" is set to -1.0. That is, a data search process is executed in which the relative search influence of "people" to a word with no weight set is set to -1.0.
[0056] Specifically, a "weighted query-related feature vector" is generated by setting the weight of "people" to -1.0, and a similarity determination process is performed between the generated "weighted query-related feature vector" and a content-related feature vector such as an image. Images with a high degree of similarity in this similarity determination process are output as search results. As a result, search results with a low influence of "people" can be obtained.
[0057] (2-(b) Configuration of a query input area that allows weight setting by bracket input and query input examples) Next, we will explain the configuration of a query input area that allows weight setting by bracket input and query input examples.
[0058] FIG. 3, like FIG. 2 described above, is a diagram illustrating the configuration of a query input area in a data search UI displayed on the display unit of an information processing device of the present disclosure and an example of query input.
[0059] Figure 3 shows examples of input of three types of queries (search statements): (1) a normal search query, (b1) Example 1 of a search query with weight set using brackets (parentheses), and (b2) Example 2 of a search query with weight set using brackets (parentheses).
[0060] The example shown in "(1) Normal Search Query" is an example in which a conventional query, i.e., the following query, is entered in the query input area 11 in the data search UI, as described above with reference to FIG. 1 .
[0061] “(b1) Example 1 of a weighted search query using bracket input” is an example of the configuration of the query input area 11b in the data search UI displayed on the display unit of the information processing device disclosed herein and an example of a query to be input into the query input area 11b.
[0062] In the query input area 11b, a query can be input in which a "weight" is set for each constituent word of the query. For example, the following weight-set query can be input: Query="(people) running on the beach" This is an example of an input query.
[0063] "(people)" means that the weight of one word that makes up the query, "people," is set to 1.1 times. Brackets (parentheses) mean that the weight of the word in the brackets () is set to 1.1 times. Note that double brackets (()) means that the weight of the word in the double brackets (()) is set to (1.1). 2 The triple brackets ((())) mean that the weight of the word inside the triple brackets ((())) is set to (1.1). 3 This means setting it to double.
[0064] The query shown in FIG. 3(b1) includes five words: "people," "running," "on," "the," and "beach."
[0065] Query="(people) running on the beach" This weight setting query is a query in which the weight of "people" is multiplied by 1.1, and the weights of "running," "on," "the," and "beach" are multiplied by 1.0.
[0066] When this weight-setting query is input, the information processing device of the present disclosure executes a data search process in which the weight of "people" is set to 1.1 times that of the other four words, "running," "on," "the," and "beach." That is, the information processing device executes a data search process in which the influence of "people" is set to 1.1 times the influence of the other four words.
[0067] Specifically, the information processing device of the present disclosure generates a "weighted query-related feature vector" in which the influence of "people" is set to 1.1 times the influence of the other four words, and executes a similarity determination process between the generated "weighted query-related feature vector" with a high influence of "people" and a content-related feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, search results with a high influence of "people" are obtained.
[0068] "(b2) Weighted Search Query Example 2 Using Bracket Input" is another example of a query to be input into the query input field 11b in the data search UI displayed on the display unit of the information processing device of the present disclosure. That is, this shows an input example of the query: Query="(people) running on the ((beach))."
[0069] "(people)" means that the weight of one word that makes up the query, "people," is set to 1.1 times. "((beach))" means that the weight of one word that makes up the query, "beach," is set to (1.1). 2 This means setting it to double.
[0070] When this weight setting query is input, the information processing device of the present disclosure increases the weight of "people" by 1.1 times and the weight of "beach" by (1.1). 2 In other words, the relative search influence of the three unweighted words (running on the) is set to 1.1 times for "people" and 1.1 times for "beach." 2 Execute the data search process set to double.
[0071] Specifically, the information processing device of the present disclosure increases the influence of "people" by 1.1 times and the influence of "beach" by 1.1 times compared to three words (running on the) that have no weight setting. 2A "weighted query-corresponding feature vector" is generated with a weighted weight multiplied by 1, and a similarity determination process is performed between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, search results with high influence of "people" and "beach" can be obtained.
[0072] The example of the weight setting query described with reference to FIG. 3 is an example in which brackets indicating weights are associated with each word and input. The brackets set for each word are within a specified range, for example, from one bracket to three brackets. However, the maximum number of brackets allowed may be within a predetermined range, or may be freely set by the user.
[0073] (2-(c) Configuration of a query input area that allows weight setting using a scroll bar and an example of query input) Next, a configuration of a query input area that allows weight setting using a scroll bar and an example of query input will be described.
[0074] 2 and 3 described above, Fig. 4 is a diagram illustrating the configuration of a query input area in a data search UI displayed on the display unit of an information processing device according to the present disclosure, and an example of query input. Fig. 4 shows two types of query (search statement) input examples: (1) a normal search query, and (c1) Example 1 of a weighted search query using a scroll bar.
[0075] The example shown in "(1) Normal Search Query" is an example in which a conventional query, i.e., the following query, is entered in the query input area 11 in the data search UI, as described above with reference to FIG. 1 .
[0076] "(c1) Example 1 of a weight-setting search query using a scroll bar" is a diagram illustrating the configuration of the query input area 11c, which allows input of a weight-setting search query using a scroll bar that can be used in the information processing device of the present disclosure, and an example of query input.
[0077] As shown in the figure, the query input area 11c has a query input section 21 and a word weight setting section 22. In the word weight setting section 22, a weight setting word input section 23 and a weight adjustment scroll bar 24 are displayed.
[0078] A conventional query, for example, the following query is input to the query input unit 21: Query="people running on the beach".
[0079] The weight setting word input section 23 of the word weight setting section 22 allows the user to input a word to be adjusted in weight selected by the user from the words constituting the query input to the query input section 21. The figure shows an example in which "people" is input.
[0080] The weight adjustment scroll bar 24 of the word weight setting unit 22 is a scroll bar for setting the weight of a word inputted to the weight setting word input unit 23. By operating the scroll bar, the user can set the weight of the word inputted to the weight setting word input unit 23 ("people" in this example).
[0081] In the example shown in the figure, the weight setting range using the weight adjustment scroll bar 24 is -1.0 to 2.0. However, this is just one example, and various weight setting ranges are possible using the weight adjustment scroll bar 24. Also, the range of weights that can be freely set by the user may be adjustable.
[0082] In the example shown in the figure, the weight adjustment scroll bar 24 is set to the weight=1.5 position, in which case the weight of "people" is set to 1.5 times.
[0083] When this weight-setting query is input, the information processing device of the present disclosure executes a data search process in which the weight of "people" is set to 1.5 times that of the other four words, "running," "on," "the," and "beach." That is, the information processing device executes a data search process in which the influence of "people" is set to 1.5 times the influence of the other four words.
[0084] Specifically, the information processing device of the present disclosure generates a "weighted query-related feature vector" in which the influence of "people" is set to 1.5 times the influence of the other four words, and executes a similarity determination process between the generated "weighted query-related feature vector" with a high influence of "people" and a content-related feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, search results with a high influence of "people" are obtained.
[0085] Fig. 5 shows "(c2) Example 2 of a weight-set search query using a scroll bar." The query input area 11c shown in Fig. 55(c2) has a query input section 21 and two word weight setting sections 22-1 and 22-2. The upper word weight setting section 22-1 has a weight-set word input section 23-1 and a weight adjustment scroll bar 24-1. The lower word weight setting section 22-2 also has a weight-set word input section 23-2 and a weight adjustment scroll bar 24-2.
[0086] A conventional query, for example, the following query is input to the query input unit 21: Query="people running on the beach".
[0087] The upper word weight setting section 22-1 and the lower word weight setting section 22-2 are areas for setting the weights of different words. In the example shown in the figure, "people" is input into the weight setting word input section 23-1 of the upper word weight setting section 22-1, and "running" is input into the weight setting word input section 23-2 of the lower word weight setting section 22-2.
[0088] The weight adjustment scroll bar 24-1 in the upper word weight setting section 22-1 is set to a weight of 1.5 times. On the other hand, the weight adjustment scroll bar 24-2 in the lower word weight setting section 22-2 is set to a weight of -0.5 times. In this case, the weight of "people" is set to 1.5 times, and the weight of "running" is set to -0.5 times.
[0089] When such a weight-setting query is input, the information processing device of the present disclosure executes a data search process in which the weight of "people" is set to 1.5 times that of the other three words "on," "the," and "beach," and the weight of "running" is set to -0.5 times that of the other three words "on," "the," and "beach."
[0090] That is, a data search process is executed in which the influence of "people" is set to 1.5 times the influence of the three unweighted words, and the influence of "running" is set to -0.5 times the influence of the three unweighted words.
[0091] Specifically, the information processing device of the present disclosure generates a "weighted query-corresponding feature vector" in which the influence of "people" is set to 1.5 times the influence of the three unweighted words and the influence of "running" is set to -0.5 times the influence of the three unweighted words, and performs a similarity determination process between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector such as an image.
[0092] In this similarity determination process, images with high similarity are output as search results, resulting in search results in which the influence of "people" is high and the influence of "running" is low.
[0093] 5 shows an example in which two word weight setting units 22 are provided, it is possible to sequentially add and display more word weight setting units 22 up to the number equal to the number of words constituting the query input in the query input area 11 c. In other words, the user can set individual weights for all of the words constituting the query input in the query input area 11 c.
[0094] (2-(d) Configuration of a query input area that allows weight setting using color display (heat map) and example query input) Next, we will explain the configuration of a query input area that allows weight setting using color display (heat map) and example query input.
[0095] 6, similar to FIGS. 2 to 5 described above, is a diagram illustrating the configuration of a query input area in a data search UI displayed on the display unit of an information processing device of the present disclosure and an example of query input.
[0096] FIG. 6 shows examples of input of two types of queries (search statements): (1) a normal search query; and (d1) weight-setting search query example 1 using color display (heat map).
[0097] The example shown in "(1) Normal Search Query" is an example in which a conventional query, i.e., the following query, is entered in the query input area 11 in the data search UI, as described above with reference to FIG. 1 .
[0098] “(d1) Weighted search query example 1 using color display (heat map)” is an example of the configuration of the query input area 11d in the data search UI displayed on the display unit of the information processing device disclosed herein and an example of a query to be entered in the query input area 11d.
[0099] In the query input field 11d, a query can be input in which a "weight" is set for each constituent word of the query. The figure shows an example in which the following query is input: Query="people running on the beach" This is an input example of the above query.
[0100] Furthermore, a dark color panel is superimposed on the word area of "people," a light color panel is superimposed on the word area of "running," and a medium color panel is superimposed on the word area of "beach."
[0101] These color panels are set by the user, and as shown in the table on the left side of FIG. 6 , each color of the panel corresponds to a different weight. As shown in the table on the left side of FIG. 6 , when the panel color is light, the weight is 1.0, and as the panel color becomes darker, the weight value increases, with the darkest panel having a weight of 3.0. Note that the correspondence data between color panels and weights shown in this figure is an example, and various other settings are possible. Alternatively, the user may be allowed to freely set the correspondence between color panels and weights.
[0102] In the example shown in Figure 6, a dark color panel (weight = 3 times) is superimposed on the word region of "people", so the weight of "people" is 3 times. A light color panel (weight = 1.5 times) is superimposed on the word region of "running", so the weight of "running" is 1.5 times. A medium color panel (weight = 2 times) is superimposed on the word region of "beach", so the weight of "beach" is 2 times.
[0103] When this weight-set query is input, the information processing device of the present disclosure executes a data search process in which the weights of "people" are set to 3 times, "running" to 1.5 times, and "beach" to 2 times. Note that these weights are relative to the unweighted words "on" and "the."
[0104] Specifically, the information processing device of the present disclosure generates a "weighted query-corresponding feature vector" by setting the relative influence (weight) of unweighted words to 3 times for "people," 1.5 times for "running," and 2 times for "beach," and then performs a similarity determination process between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector such as an image.
[0105] In this similarity determination process, images with high similarity are output as search results, resulting in search results with a high influence of "people," "running," and "beach."
[0106] 2 to 6, several configuration examples of the query input area in the data search UI (user interface) displayed on the display unit of the information processing device of the present disclosure have been described. That is, the following configurations: (a) A configuration of the query input area that allows weight setting by inputting numerical values (b) An example configuration of the query input area that allows weight setting by inputting brackets (c) A configuration of the query input area that allows weight setting using a scroll bar (d) A configuration of the query input area that allows weight setting using a color display (heat map)
[0107] The information processing device of the present disclosure displays at least one of the query input areas (a) to (d) above on a data search UI (user interface). Note that a configuration may be adopted in which multiple query input areas (a) to (d) above are displayed together, allowing the user to select and use the query input area that is easiest to use.
[0108] 3. Specific Example of Data Search Processing Using Weight Setting Query Next, a specific example of data search processing using a weight setting query will be described.
[0109] As described with reference to Figures 2 to 6, the information processing device of the present disclosure displays a data search UI (user interface) on a display unit, which has a query input area in which a weight-setting query can be input. A user can set various weights for each word constituting a query and perform a data search. A specific example of a data search process performed using a weight-setting query will be described below.
[0110] 7 and subsequent figures, the following specific examples of each process will be described in order: (1) A specific example of a data search process using a weight setting query by inputting numerical values (2) A specific example of a data search process using a weight setting query using a scroll bar (3) A specific example of a data search process using a query by Japanese text (4) A specific example of a data search process in which music data is the search target
[0111] (3-(1) Specific Example of Data Search Processing Using a Weight Setting Query Based on Numerical Input) First, a specific example of data search processing using a weight setting query based on numerical input will be described.
[0112] FIG. 7 is a diagram illustrating a specific example of data search processing using "(a) query input area 11a that allows weight setting by inputting numerical values" described above with reference to FIG.
[0113] FIG. 7 shows an example of data search processing using the following query, i.e., Query=A photo of a smiling face. Three different weight-set queries are generated by setting three different weights to one word, “smiling,” that makes up the above query.
[0114] Specifically, the following three types of weight setting queries are used to perform data search: (1) A query in which the weight of "smiling" is set to 1.0; (2) A query in which the weight of "smiling" is set to 0.4; and (3) A query in which the weight of "smiling" is set to -0.2.
[0115] In any of the above (1) to (3), the full text of the query initially entered in the query input area 11a is: Query=A photo of a smiling face.
[0116] 7(1) shows an example in which the user inputs the above query and then inputs a weighting value of 1.0 for the word "smiling" included in the query, changing the query to the following weight-setting query: Weight-setting query = A photo of a (smiling: 1.0) face This query change process sets the weight of the word "smiling" to 1.0.
[0117] 7(2) shows an example in which the user inputs the above query and then inputs a weighting value of 0.4 for the word "smiling" included in the query, thereby changing the query to the following weight-setting query: Weight-setting query = A photo of a (smiling: 0.4) face This query change process sets the weight of the word "smiling" to 0.4 times.
[0118] 7(3) shows an example in which the user inputs the above query and then inputs a weighting value of −0.2 for the word “smiling” included in the query, changing the query to the following weight-setting query: Weight-setting query = A photo of a (smiling: −0.2) face This query change processing sets the weight of the word “smiling” to −0.2 times.
[0119] The right side of Fig. 7 shows image data as the results of a search performed by the data search process using each of the weight setting queries (1) to (3). These search results are displayed in the search result display area 13 of the data search UI displayed on the display unit of the information processing device described above with reference to Fig. 1.
[0120] 7 upper part = A photo of a (smiling: 1.0) face In the data search process using this weight setting query, the search process is executed with the weight of the word "smiling" set to 1.0. As a result, "smiling", i.e., face image data with a large element of laughter, is output as the search result.
[0121] 7 shows a weight setting query (2) = A photo of a (smiling: 0.4) face. In the data search process using this weight setting query, the weight of the word "smiling" is set to 0.4. As a result, "smiling," i.e., face image data with a low level of laughter, is output as the search result.
[0122] 7, the weight setting query (3) = A photo of a (smiling: -0.2) face. In the data search process using this weight setting query, the search process is executed with the weight of the word "smiling" set to -0.2. As a result, "smiling," i.e., face image data with little laughing element, is output as the search result.
[0123] As described above, the information processing device of the present disclosure generates a "weight-set query-related feature vector" based on a weight-set query, and executes a similarity determination process between the generated "weight-set query-related feature vector" and a content-related feature vector of an image, etc. Images with high similarity in this similarity determination process are output as search results.
[0124] As a result, a search process is executed that reflects the weight of each word set in the weight-setting query. In the example shown in Figure 7, the weight of one word "smiling" in the weight-setting query is set to three values: 1.0, 0.4, and -0.2, and the results of a data search are shown.
[0125] As can be seen from the data search results shown in Figure 7, it is possible to perform a data search that selects and acquires images according to the weight set by the user for the word "smiling" in the query, i.e., images with different levels of laughter.
[0126] FIG. 8 is also a diagram illustrating a specific example of data search processing using "(a) query input area 11a that allows weight setting by inputting numerical values" previously described with reference to FIG.
[0127] 8 shows an example of data search processing using three different weight-set queries generated by setting different weights to the two words "sunny" and "foggy" that make up the following query: Query = Trees stand quietly with flowers. It is sunny and foggy.
[0128] Specifically, the following three types of weighted queries are used to perform data searches: (1) A query with no weights set (i.e., a query in which the weights of "sunny" and "foggy" are set to 1.0) (2) A query in which the weight of "sunny" is set to 1.5 and the weight of "foggy" is set to -0.5 (3) A query in which the weight of "sunny" is set to 1.2 and the weight of "foggy" is set to 2.0
[0129] In any of the above (1) to (3), the full text of the query initially entered in the query input area 11a is: Query=Trees stand quietly with flowers. It is sunny and foggy. This is the above query.
[0130] FIG. 8(1) shows an example in which the user inputs the above query and then executes a search without setting weights for the words included in the query.
[0131] 8(2) shows an example in which, after inputting the above query, the user inputs a weighting value of 1.5 for the word "sunny" included in the query, and further inputs a weighting value of -0.5 for the word "foggy", thereby changing the query to the following weight-setting query: Weight-setting query = Trees stand quietly with flowers. It is (sunny: 1.5) and (foggy: -0.5). This query change process sets the weight of the word "sunny" to 1.5 times and the weight of the word "foggy" to -0.5 times.
[0132] 8(3) shows an example in which, after inputting the above query, the user inputs a weighting value of 1.2 for the word "sunny" included in the query, and further inputs a weighting value of 2.0 for the word "foggy," thereby changing the query to the following weight-setting query: Weight-setting query = Trees stand quietly with flowers. It is (sunny: 1.2) and (foggy: 2.0). This query change process sets the weight of the word "sunny" to 1.2 times, and the weight of the word "foggy" to 2.0 times.
[0133] The right side of Fig. 8 shows image data as the results of a search performed by a data search process using each of the weight setting queries (1) to (3). These search results are displayed in the search result display area 13 of the data search UI displayed on the display unit of the information processing device described above with reference to Fig. 1.
[0134] Query (1) shown in the upper part of Figure 8, i.e., Query = Trees stand quietly with flowers. It is sunny and foggy. In a data search process using this unweighted query, the search process is executed with the same weight (1.0) set for all words constituting the query. As a result, image data in which the words "sunny" and "foggy" are reflected to the same extent is output as the search result.
[0135] Furthermore, the weight setting query (2) shown in the middle of FIG. 8, i.e., Weight setting query = Trees stand quietly with flowers. It is (sunny: 1.5) and (foggy: -0.5). In the data search process using this weight setting query, the search process is executed with the weight of the word "sunny" set to 1.5 times and the weight of the word "foggy" set to -0.5 times. As a result of this search process, image data with a high "sunny" (i.e., sunshine element) and a low "foggy" (i.e., foggy element) is output as the search result.
[0136] Furthermore, the weight setting query (3) shown in the lower part of Figure 8, i.e., Weight setting query = Trees stand quietly with flowers. It is (sunny: 1.2) and (foggy: 2.0). In the data search process using this weight setting query, the search process is executed with the weight of the word "sunny" set to 1.2 times and the weight of the word "foggy" set to 2.0 times. As a result of this search process, image data with a low "sunny" (sunlight) element and a high "foggy" (foggy) element is output as the search result.
[0137] As described above, the information processing device of the present disclosure generates a "weight-set query-related feature vector" based on a weight-set query, and executes a similarity determination process between the generated "weight-set query-related feature vector" and a content-related feature vector such as an image. Images with high similarity in this similarity determination process are output as search results. As a result, a search process is executed that reflects the weight of each word set in the weight-set query.
[0138] In the example shown in Figure 8 (2), a "weighted query-related feature vector" is generated based on weight setting information in which the weight of the word "sunny" is set to 1.5 times and the weight of the word "foggy" is set to -0.5 times, and a similarity determination process is performed between the generated "weighted query-related feature vector" and a content-related feature vector of an image, etc. As a result, image data that is high in "sunny," i.e., the element of sunlight, and low in "foggy," i.e., the element of foggyness, can be obtained as a search result.
[0139] In the example shown in FIG. 8(3), a "weighted query-corresponding feature vector" is generated based on weight setting information in which the weight of the word "sunny" is multiplied by 1.2 and the weight of the word "foggy" is multiplied by 2.0, and a similarity determination process is performed between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector such as an image. Images with a high similarity in this similarity determination process are output as search results. As a result, image data with a low "sunny" (i.e., sunshine) element and a high "foggy" (i.e., foggy) element is output as search results.
[0140] (3-(2) Specific Example of Data Search Processing Using a Weight Setting Query with a Scroll Bar) Next, a specific example of data search processing using a weight setting query with a scroll bar will be described.
[0141] FIG. 9 is a diagram illustrating a specific example of a data search process using “(c) a query input area 11c that allows weight setting using a scroll bar,” which was previously described with reference to FIGS. 4 and 5.
[0142] 9, the query input area 11c has a query input section 21 and a word weight setting section 22. In the word weight setting section 22, a weight setting word input section 23 and a weight adjustment scroll bar 24 are displayed.
[0143] The following query, that is, Query=A bear with a ribbon 3D style painting, is input into the query input section 21 of the query input area 11c.
[0144] The figure also shows a state in which "3D" is input into the weighted word input section 23 of the word weight setting section 22, and the user is adjusting the weight adjustment scroll bar 24 of the word weight setting section 22. In the example shown, the weight setting range using the weight adjustment scroll bar 24 is -1.0 to 2.0. By moving the weight adjustment scroll bar 24 left and right, the user can adjust the weight of "3D" input into the weighted word input section 23 of the word weight setting section 22 within the range of -1.0 to 2.0.
[0145] 9 shows an example of search results that are output by moving the weight adjustment scroll bar 24 left or right. Note that the actual search results will be a combination of various bear images, but for the sake of explanation, the image of one type of bear will be used as a sample.
[0146] For example, when the user moves the weight adjustment scroll bar 24 to the left end, a data search process is executed in which the weight of "3D" input in the weight setting word input section 23 of the word weight setting section 22 is set to -1.0.
[0147] That is, a data search process is executed for the query entered in the query input section 21 of the query input area 11c, i.e., "Query = A bear with a ribbon 3D style painting." The weight of the word "3D" in the query is set to -1.0. As a result, a bear image with few 3D elements, i.e., a nearly 2D image of a bear, as shown in the lower left corner of Figure 9, is output as a search result.
[0148] On the other hand, when the user moves the weight adjustment scroll bar 24 to the right end, a data search process is executed in which the weight of "3D" input into the weight setting word input section 23 of the word weight setting section 22 is set to 2.0 times.
[0149] That is, a data search process is executed for the query entered in the query input section 21 of the query input area 11c, i.e., "Query = A bear with a ribbon 3D style painting." The weight of the word "3D" in the query is set to 2.0. As a result, a bear image with many 3D elements, i.e., a nearly 3D image of a bear, as shown in the lower right corner of Figure 9, is output as a search result.
[0150] By moving the weight adjustment scroll bar 24 left and right, the user can adjust the weight of "3D" entered in the weight setting word input section 23 of the word weight setting section 22 within the range of -1.0 to 2.0, making it possible to obtain image data having 3D elements corresponding to each weight as search results.
[0151] In this example, too, the information processing device of the present disclosure generates a "weight-setting query-corresponding feature vector" based on the weight-setting query, and performs a similarity determination process between the generated "weight-setting query-corresponding feature vector" and a content-corresponding feature vector such as an image.
[0152] 9, by changing the weight of the word "3D" between -1.0 and 2.0, a "weighted query-corresponding feature vector" corresponding to each weight is generated, and a similarity determination process is performed between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector such as an image, and images with high similarity are output as search results. As a result, images with different 3D levels according to the weight level of "3D," i.e., 2D to 3D image data, can be obtained as search results.
[0153] Like FIG. 9, FIG. 10 is a diagram illustrating a specific example of a data search process using "(c) a query input area 11c that allows weight setting using a scroll bar" as previously described with reference to FIGS. 4 and 5.
[0154] 10, the query input area 11c has a query input section 21 and a word weight setting section 22. In the word weight setting section 22, a weight setting word input section 23 and a weight adjustment scroll bar 24 are displayed.
[0155] The following query, that is, Query=A photo of crowded spaces, is input into the query input section 21 of the query input area 11c.
[0156] The figure also shows a state in which "crowdded" is input into the weight setting word input section 23 of the word weight setting section 22, and the user is adjusting the weight adjustment scroll bar 24 of the word weight setting section 22. In the example shown, the weight setting range using the weight adjustment scroll bar 24 is -1.0 to 2.0 times. By moving the weight adjustment scroll bar 24 left and right, the user can adjust the weight of "crowdded" input into the weight setting word input section 23 of the word weight setting section 22 within the range of -1.0 to 2.0 times.
[0157] 10 shows an example of search results that are output by moving the weight adjustment scroll bar 24 left or right. For example, when the user moves the weight adjustment scroll bar 24 to the left end, a data search process is executed in which the weight of "crowded" input in the weight setting word input section 23 of the word weight setting section 22 is set to -1.0.
[0158] That is, a data search process is executed in which the weight of the word "crowded" in the query input section 21 of the query input area 11c, i.e., "Query = A photo of crowded spaces," is set to -1.0. As a result, image data with a low level of congestion, as shown in the lower left corner of FIG. 10, is output as a search result.
[0159] On the other hand, when the user moves the weight adjustment scroll bar 24 to the right end, a data search process is executed in which the weight of "crowded" input in the weight setting word input section 23 of the word weight setting section 22 is set to 2.0 times.
[0160] That is, a data search process is executed in which the weight of the word "crowded" in the query input section 21 of the query input area 11c, i.e., "Query = A photo of crowded spaces," is set to 2.0. As a result, image data with a high level of congestion, as shown in the lower right corner of FIG. 10, is output as a search result.
[0161] By moving the weight adjustment scroll bar 24 left and right, the user can adjust the weight of "crowded" entered in the weight setting word input section 23 of the word weight setting section 22 within the range of -1.0 to 2.0, making it possible to obtain image data of the congestion level corresponding to each weight as a search result.
[0162] In this example, too, the information processing device of the present disclosure generates a "weight-setting query-corresponding feature vector" based on the weight-setting query, and performs a similarity determination process between the generated "weight-setting query-corresponding feature vector" and a content-corresponding feature vector such as an image.
[0163] In the example shown in Fig. 10, by changing the weight of the word "crowded" between -1.0 and 2.0 times, a "weighted query-corresponding feature vector" corresponding to each weight is generated, and a similarity determination process is performed between the generated "weighted query-corresponding feature vector" and a content-corresponding feature vector of an image, etc., so that images with high similarity are output as search results. As a result, image data with different congestion levels according to the weight level of "crowded" can be obtained as search results.
[0164] (3-(3) Specific Example of Data Search Processing Using a Query in a Language Other Than English) Next, a specific example of data search processing using a query in a language other than English will be described.
[0165] The specific examples of search processing described with reference to Figures 7 to 10 all use English sentences as search queries, but it is also possible to generate search queries or weight setting queries using languages other than English and perform data searches.
[0166] Fig. 11 is a diagram illustrating a specific example of data search processing using a query in Japanese text. Fig. 11 is a processing example in which a search process similar to the data search processing previously described with reference to Fig. 7 is executed by applying a query generated using Japanese text, and is a diagram illustrating a specific example of data search processing using "(a) query input area 11a that allows weight setting by numerical input" previously described with reference to Fig. 2.
[0167] FIG. 11 shows an example of data search processing using three different weight-set queries generated by setting three different weights to the single word "laugh" that constitutes the following query: Query = Photo of a smiling face.
[0168] Specifically, the following three types of weighted query are used to perform data search: (1) A query in which the weight of "laugh" is set to 1.0; (2) A query in which the weight of "laugh" is set to 0.4; and (3) A query in which the weight of "laugh" is set to -0.2.
[0169] In any of the above (1) to (3), the full text of the query initially entered in the query input area 11a is: Query=Photo of a smiling face.
[0170] 11(1) shows an example in which the user inputs the above query and then inputs a weight value of 1.0 for the word "laugh" included in the query, changing it to the following weight-setting query: Weight-setting query = (laugh: 1.0) Photo of face This query change process sets the weight of the word "laugh" to 1.0.
[0171] 11(2) shows an example in which the user inputs the above query and then inputs a weighting value of 0.4 for the word "laugh" included in the query, changing it to the following weight-setting query: Weight-setting query = (laugh: 0.4) Photo of face This query change process sets the weight of the word "laugh" to 0.4 times.
[0172] 11(3) shows an example in which the user inputs the above query and then inputs a weighting value of −0.2 for the word “laugh” included in the query, changing it to the following weight-setting query: Weight-setting query = (laugh: −0.2) Photo of face This query change processing sets the weight of the word “laugh” to −0.2 times.
[0173] The right side of Fig. 11 shows image data as a result of a search performed by a data search process using each of the weight setting queries (1) to (3). These search results are displayed in the search result display area 13 of the data search UI displayed on the display unit of the information processing device described above with reference to Fig. 1.
[0174] 11 , the weighted query (1) shown in the upper part of the figure = (laughing: 1.0) face photo In the data search process using this weighted query, the search process is executed with the weight of the word "laugh" set to 1.0. As a result, "laughing," i.e., face image data with a large element of laughter, is output as the search result.
[0175] 11 shows a weighted query (2) = (laughing: 0.4) face photo. In the data search process using this weighted query, the weight of the word "laughing" is set to 0.4. As a result, "laughing," i.e., face image data with a low laughing element, is output as the search result.
[0176] Furthermore, in the data search process using the weighted query (3) shown in the lower part of Fig. 11 = (laughing: -0.2) face photo, the weight of the word "laughing" is set to -0.2. As a result, "laughing," i.e., face image data with few laughing elements, is output as the search result.
[0177] As described above, the information processing device of the present disclosure generates a "weight-set query-related feature vector" based on a weight-set query, and executes a similarity determination process between the generated "weight-set query-related feature vector" and a content-related feature vector of an image, etc. Images with high similarity in this similarity determination process are output as search results.
[0178] As a result, a search process is executed that reflects the weight of each word set in the weight-setting query. In the example shown in Figure 11, the weight of one word in the weight-setting query, "laugh," is set to three values: 1.0, 0.4, and -0.2, and the results of a data search are shown.
[0179] The data search results in each of Figures 11(1) to (3) will output images according to the weighting set by the user for the word "laugh" in the query, i.e., images with different levels of laughter.
[0180] (3-(4) Specific Example of Data Search Processing in Which Music Data is the Search Target) Next, a specific example of data search processing in which music data is the search target will be described.
[0181] In the above specific example, an example of search processing using image data as the data search target has been described, but the information processing device of the present disclosure is not limited to searching for data in categories other than images. For example, it is possible to perform a data search using data in various categories, such as documents, music, architecture, nature, and living things, in addition to images.
[0182] An example of search processing when music is set as the search target will be described with reference to FIG.
[0183] FIG. 12 is a specific example of a data search process using “(a) a query input area 11 a that allows weighting by numerical input” as described above with reference to FIG. 2, and is a diagram illustrating an example of a search process in which music is the search target.
[0184] FIG. 12 shows an example of data search processing using three different weight-set queries generated by setting three different weights to the two words that make up the query: "bright" and "lively."
[0185] Specifically, the following three types of weighted query are used to provide examples of data searches: (1) A query in which the weight of "bright" is set to 1.5 times; (2) A query in which the weight of "bright" is set to 2.0 times and the weight of "lively" is set to 2.0 times; (3) A query in which the weight of "bright" is set to 1.0 times and the weight of "lively" is set to -1.0 times.
[0186] In any of the above (1) to (3), the full query initially entered in the query input area 11a is: Query = cheerful, lively music.
[0187] 12(1) shows an example in which the user inputs the above query and then inputs a weight value of 1.5 for the word "bright" included in the query, changing it to the following weight-setting query: Weight-setting query = (bright: 1.5), lively music This query change process increases the weight of the word "bright" by 1.5 times.
[0188] 12(2) shows an example in which, after inputting the above query, the user inputs a weighting value of 2.0 for the word "bright" included in the query, and further inputs a weighting value of 2.0 for the word "lively", thereby changing the query to the following weight-setting query: Weight-setting query = (bright: 2.0), (lively: 2.0) music This query change process sets the weight of the word "bright" to 2.0 times, and also sets the weight of the word "lively" to 2.0 times.
[0189] 12(3) shows an example in which, after inputting the above query, the user inputs a weighting value of 1.2 for the word "bright" included in the query, and further inputs a weighting value of -1.0 for the word "lively", thereby changing the query to the following weight-setting query: Weight setting query = (bright: 1.2), (lively: -1.0) music This query change process sets the weight of the word "bright" to 1.2 times, and further sets the weight of the word "lively" to -1.0 times.
[0190] The right side of Fig. 12 shows image data as the results of a search performed by the data search process using each of the weight setting queries (1) to (3). These search results are displayed in the search result display area 13 of the data search UI displayed on the display unit of the information processing device described above with reference to Fig. 1.
[0191] 12, the weighted query (1) = (bright: 1.5), lively music. In the data search process using this weighted query, the weight of the word "bright" is set to 1.5 times. As a result, music data with a large amount of bright elements is output as the search result.
[0192] 12 shows a weighted query (2) = (bright: 2.0), (lively: 2.0) music. In the data search process using this weighted query, the weight of the word "bright" is set to 2.0 and the weight of the word "lively" is also set to 2.0. As a result, music data with large amounts of both bright and lively elements is output as the search result.
[0193] Furthermore, in the weighted query (3) shown in the lower part of Fig. 12, music is (bright: 1.2), (lively: -1.0). In the data search process using this weighted query, the weight of the word "bright" is set to 1.2 times, and the weight of the word "lively" is set to -1.0. As a result, music data with many bright elements but few lively elements is output as the search result.
[0194] As described above, the information processing device of the present disclosure generates a "weight-set query-related feature vector" based on a weight-set query, and executes a similarity determination process between the generated "weight-set query-related feature vector" and a content-related feature vector of an image, etc. Images with high similarity in this similarity determination process are output as search results.
[0195] As a result, a search process is executed that reflects the weight of each word set in the weight-setting query. The example shown in Figure 12 shows the results of a data search performed by changing the weight of one word in the weight-setting query, "bright" and the weight of "lively."
[0196] The data search results for each of (1) to (3) in Figure 12 will output music according to the weights set by the user for the words "bright" and "lively" in the query, i.e., music with different levels of brightness and liveliness.
[0197] 4. Details of the Configuration and Processing of the Information Processing Apparatus of the Present Disclosure Next, details of the configuration and processing of the information processing apparatus of the present disclosure will be described.
[0198] The information processing device of the present disclosure is a device capable of performing data search processing, such as a PC, a smartphone, a tablet terminal, etc. Note that the information processing device of the present disclosure can be configured as a single device, but can also be configured as multiple devices.
[0199] Fig. 13 shows an example configuration of an information processing device 100 according to the present disclosure. Note that the configuration diagram shown in Fig. 13 does not show the overall configuration of an information processing device such as a PC, but is a configuration diagram showing the main components used in the processing according to the present disclosure. The configuration of the information processing device 100 shown in Fig. 13 will be described. As shown in Fig. 13, the information processing device 100 has a UI (user interface) 101, a data processing unit 102, and a communication unit 103.
[0200] The UI (user interface) 101 has an input unit 111 and an output unit 112. The output unit 112 is configured by, for example, a display unit that outputs display data for the data search UI, etc., previously described with reference to Fig. 1. The input unit 111 is configured by at least one input means, such as a touch panel, a keyboard, a mouse, etc.
[0201] A user can input a query (search statement) and weight information for each word constituting the query via the input unit 111 of the UI (user interface) 101. The user can input a query (search statement) and weight information for each word constituting the query by using the various types of query input areas 11a to 11c previously described with reference to FIGS. 2 to 12. As shown in FIG. 13, this input information is input from the input unit 111 to the data processing unit 102. The query (search statement) and weight information shown in FIG. 13 are shown.
[0202] The data processing unit 102 receives a query (search statement) and weight information for each word that constitutes the query via the input unit 111 of the UI (user interface) 101, and executes a data search process using this input information. The data to be searched is, for example, data stored in databases 200-1 to 200-n that can be accessed via the communication unit 103, data on the Internet, etc.
[0203] The data processing unit 102 executes a search process for content such as images that can be accessed via the communication unit 102 based on the query entered by the user and the weight information. The search results generated in this search process are output to the output unit of the UI (user interface) 101. These search results are shown in Fig. 13. These search results are displayed, for example, in the search result display area 13 of the data search UI described above with reference to Fig. 1.
[0204] When executing a data search process, the data processing unit 102 generates a "weighted query-corresponding feature vector" based on the query and weight information input by the user, and executes a similarity determination process between the generated "weighted query-corresponding feature vector" and a feature vector generated corresponding to content such as an image accessible via the communication unit 102. Images with high similarity in this similarity determination process are set as search results.
[0205] The detailed configuration of the data processing unit 102 and the details of the processing to be executed will be described with reference to Fig. 14. As shown in Fig. 14, the data processing unit 102 includes a text encoder 121, a weighted query-related feature vector generation unit 122, and a similarity analysis unit 123.
[0206] The text encoder 121 receives a query (search sentence) input via the input unit 111 and performs encoding processing using a learning model. A specific example of the text encoder 121 is a Transformer, which is a deep learning model (Deep Neural Network) that uses an attention mechanism that analyzes the association between words.
[0207] An overview of a deep learning model using an attention mechanism that can be used as the text encoder 121 will be described with reference to Fig. 15. The attention mechanism is a mechanism that analyzes the association between each word that makes up a query and other words. Fig. 15 shows the following two examples of attention mechanisms: (a) Unidirectional Attention (Transformer) (b) Bidirectional Attention (BERT)
[0208] "(a) Unidirectional Attention" analyzes the relevance of each word in a query to other words by analyzing only the relevance of each word in the query (search sentence) to the words following that word. In other words, it analyzes the relevance between words in only one direction, from forward to backward. This "(a) Unidirectional Attention" is used as a Transformer, which is a deep learning model (Deep Neural Network).
[0209] In contrast, "(b) bidirectional attention" analyzes the relevance of each word in a query to other words by analyzing not only the subsequent words but also the preceding words for each word in the query (search sentence). In other words, it analyzes the bidirectional inter-word relevance by analyzing not only the relevance between preceding words and subsequent words, but also the relevance between subsequent words and preceding words.
[0210] This "(b) bidirectional attention" is used as BERT (Bidirectional Encoder Representations from Transformer), which is a derivative of Transformer.
[0211] In the example shown in FIG. 15, both (a) and (b) show processing examples in which the following query, i.e., Query (search statement)=dog running, is input as an example of the query (search statement) to be analyzed.
[0212] Note that [SOS] indicates Start of sentence, [EOS] indicates End of sentence, and [PAD] indicates padding data for adjusting data length. In both "(a) Unidirectional Attention" and "(b) Bidirectional Attention," the five elements [SOS], dog, running, [EOS], and [PAD] are used as input elements, and processing is performed individually for each of the five elements from [SOS] to [PAD].
[0213] That is, each element to be processed is input to a deep learning model that uses an attention mechanism, and a "query-associated feature vector," which is a feature vector indicating the characteristics of the query (search statement), is output as an output. Note that, hereinafter, elements such as words that are input to a deep learning model with an attention mechanism and serve as processing units are referred to as "tokens." In the example shown in Figure 15, in both (a) and (b), each of the five elements [SOS], dog, running, [EOS], and [PAD] is a token.
[0214] 15 , in "(a) unidirectional attention," a "query-associated feature vector" indicating the features of the query (search sentence), which is the output, is output from the output unit of the token [EOS]. In "(a) unidirectional attention," the "query-associated feature vector" output from the output unit of the token [EOS] is generated as a feature vector including tokens such as words in the front part of the token [EOS], i.e., [SOS], dog, running, and the analysis results for these tokens, i.e., [SOS], dog, running, and [EOS], and the analysis results of token associations in one direction (from front to back) for all tokens.
[0215] On the other hand, as shown in FIG. 15, in "(b) Bidirectional Attention", first, [token corresponding vectors] output individually from the output units of all of these tokens, [SOS], dog, running, [EOS], and [PAD], are obtained, and calculation processing is performed on the multiple [token corresponding vectors] corresponding to all of these tokens, such as calculation processing of an average vector, to output a "query corresponding feature vector" that indicates the features of the query (search statement), which is the final output.
[0216] In addition, in "(a) unidirectional attention", as in "(b) bidirectional attention", the calculation process of "token correspondence vectors" corresponding to all of these tokens, [SOS], dog, running, [EOS], and [PAD], is performed internally. In "(a) unidirectional attention", from all of these "token correspondence vectors", only the "token correspondence vector" output from the output unit of [EOS] is selected and output as the "query-corresponding feature vector".
[0217] The "query-associated feature vectors" output from both "(a) unidirectional attention" and "(b) bidirectional attention" are multidimensional vectors that represent the features of the query (search term). However, because the direction of association analysis between words differs between (a) and (b), the vectors generated are slightly different.
[0218] The text encoder 121 in the data processing unit 102 of the information processing device 100 of the present disclosure shown in FIG. 14 is configured as an encoder using "(a) unidirectional attention" shown in FIG. 15(a).
[0219] This is because the weighted query-responsive feature vector generation unit 122, which is set after the text encoder 121 shown in FIG. 14, calculates the "weighted query-responsive feature vector 131."
[0220] The weighted query-corresponding feature vector generation unit 122, which is located downstream of the text encoder 121 shown in FIG. 14, calculates the difference between adjacent tokens in the "token correspondence vector" generated by the text encoder 121, and performs processing such as multiplying the calculated difference by a weight corresponding to the word (token) set by the user to calculate a "weighted query-corresponding feature vector 131."
[0221] In "(a) Unidirectional Attention" shown in Figure 15(a), a "token correspondence vector" corresponding to each token such as a word is calculated by performing a unidirectional token association analysis, so the difference in the "token correspondence vector" between adjacent tokens is data that reflects the difference in the feature vector between each token, i.e., each word. However, in "(b) Bidirectional Attention", a "token correspondence vector" corresponding to each token such as a word is calculated by performing a bidirectional token association analysis, so the difference in the "token correspondence vector" between adjacent tokens is not data that directly reflects the difference in the feature vector between each token, i.e., each word.
[0222] For this reason, the text encoder 121 in the data processing unit 102 of the information processing device 100 of the present disclosure shown in FIG. 14 is configured as an encoder of the "(a) unidirectional attention" type shown in FIG. 15(a).
[0223] As shown in FIG. 14, a weighted query-based feature vector generation unit 122 is provided downstream of a text encoder 121 in a data processing unit 102 of an information processing device 100 according to the present disclosure.
[0224] As described above, the weighted query-corresponding feature vector generation unit 122 calculates the difference between adjacent tokens in the “token correspondence vector” generated by the text encoder 121, and performs processing such as multiplying the calculated difference by a weight corresponding to the word (token) set by the user to calculate the “weighted query-corresponding feature vector 131.”
[0225] A specific example of the processing executed by the weighted query-responsive feature vector generation unit 122 will be described with reference to Fig. 16. The diagram on the left side of Fig. 16 is a diagram showing detailed configurations of the processing executed by the text encoder 121 and the weighted query-responsive feature vector generation unit 122. The right side of Fig. 16 shows the sequence of processing executed by the weighted query-responsive feature vector generation unit 122.
[0226] 16 shows a detailed configuration of the processing executed by the text encoder 121 at the bottom left, and a detailed configuration of the processing executed by the weighted query-based feature vector generation unit 122 at the top. As explained with reference to FIG. 15, the query (search statement) to be analyzed is the following query, i.e., Query (search statement) = dog running. This shows an example of processing when the above query is input. As explained with reference to FIG. 15, each of the five elements, [SOS], dog, running, [EOS], and [PAD], is set as a token serving as a processing unit.
[0227] The text encoder 121 shown in the lower left of Fig. 16 is the "(a) unidirectional attention" type encoder described with reference to Fig. 15(a). As described above, the text encoder 121, which is a "(a) unidirectional attention" type encoder, performs a calculation process of "token correspondence vectors" corresponding to all of the tokens [SOS], dog, running, [EOS], and [PAD].
[0228] 16 shows a detailed configuration of the processing executed by the weighted query-corresponding feature vector generation unit 122. The weighted query-corresponding feature vector generation unit 122 calculates the difference between adjacent tokens in the "token correspondence vector" generated by the text encoder 121, and performs processing such as multiplying the calculated difference by a weight corresponding to the word (token) set by the user to calculate the "weighted query-corresponding feature vector 131."
[0229] The weighted query-responsive feature vector generation unit 122 sequentially executes the processes of step S01 and step S02 shown on the right side of Fig. 16 to calculate the "weighted query-responsive feature vector 131." Each of the processes of steps S01 to S02 will now be described.
[0230] (Step S01) First, in step S01, the weighted query-corresponding feature vector generation unit 122 calculates a token-corresponding vector difference (D X ) calculation process is performed.
[0231] The weighted query-corresponding feature vector generation unit 122 calculates the difference between adjacent tokens in the “token correspondence vector” generated by the text encoder 121 as a token correspondence vector difference (D X ) is calculated as the token corresponding vector difference (D X ) into the adjacent token correspondence vector (V X ~V X-1 The difference (Diff) between Dx and V is calculated according to the following formula (Formula 1): Dx = V X -V X-1 ...(Formula 1)
[0232] In the example shown on the left side of FIG. 16, the token-corresponding vector difference (D 1 ) is a token correspondence vector (V 1 ) and the token (SOS) corresponding token vector (V 0 ) and the difference, that is, D 1 =V 1 -V 0 Calculate using the above formula.
[0233] In addition, the token-corresponding vector difference (D 2 ) is a token correspondence vector (V 2 ) and a token correspondence vector (V 1 ) and the difference, that is, D 2 =V 2 -V 1The vectors corresponding to adjacent tokens (V X ~V X-1 ) are calculated in sequence.
[0234] When the number of tokens from the token (SOS) to the token (EOS) is k, the k token corresponding vector difference (D 1 ~D k ) is calculated.
[0235] In addition, each token (word) corresponding vector (V 0 ~V N ) is the token-corresponding vector difference (D X ) can be shown as follows: V 1 =V 0 +D 1 V 2 =V 1 +D 2 =V 0 +D 1 +D 2 : V K =V K-1 +D K =V 0 +D 1 +...+D K
[0236] In addition, V K is a feature vector equivalent to the query-corresponding feature vector (Vec) calculated by the encoder using the unidirectional attention-based transformer shown in FIG. 15(a).
[0237] (Step S02) Next, in step S02, the weighted query-corresponding feature vector generation unit 122 generates a weighted query-corresponding feature vector V' K " is calculated.
[0238] The weight (scale) of each word (x) input by the user is x Then, according to the following formula (Formula 2), a "weighted query-corresponding feature vector V' K " is calculated. K =V 0 +s 1 D1 +...+s K D K ...(Formula 2)
[0239] However, V 0 is the leading token correspondence vector output by the encoder, s n is the weight assigned to the n-th word in the query, D n is the difference between the token correspondence vector of the nth word and the token correspondence vector of the (n-1)th word among the constituent words of the query,
[0240] In addition, the above (Equation 2) is a process of adding weights (s) corresponding to the second and subsequent tokens (x) to the token corresponding vector of the first token in the processing target element (token). x ) and the token corresponding vector difference (D X ) and the multiplication value (s X D X ) is added together. By this calculation, the "weighted query-corresponding feature vector V' K " is calculated.
[0241] The weighted query-corresponding feature vector V′ calculated in this way is K " is a feature vector that reflects the weights set by the user on a word-by-word basis. For example, if the input query (search statement) is the following query (search statement) = dog running, and the weight of the word (dog) is 2.0 times, and the weight of the word (running) is 1.0 times, then the "weight-set query-corresponding feature vector V'" calculated according to the above (Equation 2) is K " will generate a vector in which the word (dog) has a more important meaning than the word (running).
[0242] Furthermore, for example, in the query (search statement) described above with reference to FIG. 2, that is, Query (search statement)=(people:2.0) running on the beach As described above, if the weight of the word (people) is 2.0 and the weights of the other words (running, on, the, beach) are 1.0, then the "weighted query-corresponding feature vector V'" calculated according to the above (Equation 2) isK " is generated as a vector in which the word (people) has a more important meaning than the other words (running, on, the, beach).
[0243] Returning to Fig. 14, the description of the configuration of the data processing unit 102 of the information processing device 100 of the present disclosure will continue. As described above, the weighted query-corresponding feature vector generation unit 122 calculates the difference between adjacent tokens in the "token correspondence vector" generated by the text encoder 121, and performs processing such as multiplying the calculated difference by a weight corresponding to a word (token) set by the user to calculate the "weighted query-corresponding feature vector 131". That is, the weight (scale) for each word (x) input by the user is calculated by s x Then, according to the following formula (Formula 2), a "weighted query-corresponding feature vector V' K " 131 is calculated. V' K =V 0 +s 1 D 1 +...+s K D K ...(Formula 2)
[0244] As shown in FIG. 14 , the “weighted query-related feature vector 131 ” calculated by the weighted query-related feature vector generation unit 122 is input to the feature vector similarity analysis unit 123 .
[0245] The feature vector similarity analysis unit 123 performs a similarity determination between the "weighted query-corresponding feature vector 131" calculated by the weighted query-corresponding feature vector generation unit 122 and a large number of data-corresponding feature vectors included in the search target data, i.e., content-corresponding feature vectors 132.
[0246] The search target data is, for example, online data accessible via the communication unit 103 of the information processing device 100, such as content such as image data. Various content such as image data is stored in the databases 200-1 to 200-n shown in Fig. 14. In addition to the content such as image data, the databases 200-1 to 200-n also store content identifiers (image IDs, etc.) and feature vectors (content-associated feature vectors) corresponding to the characteristics of the content (images, etc.) that have been previously analyzed, in association with each other.
[0247] An example of data stored in databases 200-1 to 200-n will be described with reference to Fig. 17. The example shown in Fig. 17 is an example of a database that stores image data. The image data shown in the upper left of Fig. 17 is data stored in database 200-x. First, the management device for database 200-x assigns an identifier (ID) to each piece of image data. Next, each image is input to content encoder 211, which calculates a content-associated feature vector (Vec.x) that indicates the characteristics of each image.
[0248] A deep learning model (Deep Neural Network) that performs image analysis and calculates feature vectors is used in the content encoder 211. Specifically, similar to the above-mentioned Transformer, a Vision Transformer (ViT), a Swin Transformer, or a ResNet that performs image analysis using an attention mechanism can be used.
[0249] The content-associated feature vector (Vecx) that indicates the features of the image calculated by the content encoder 211 is stored in the database in association with the content (image, etc.) and the content ID (image ID).
[0250] The feature vector similarity analysis unit 123 of the data processing unit 102 of the information processing device 100 shown in FIG. 14 acquires content-associated feature vectors 132-1 to 132-n of each content (image) from the databases 200-1 to 200-n on the Internet via the communication unit 103.
[0251] The feature vector similarity analysis unit 123 analyzes the similarity of these content-corresponding feature vectors 132-1 to 132-n with the "weighted query-corresponding feature vector 131" calculated by the weighted query-corresponding feature vector generation unit 122. The similarity analysis process between feature vectors is performed by calculating an inter-vector similarity index value such as cosine similarity or inter-vector distance (Euclidean distance).
[0252] The feature vector similarity analysis unit 123 selects content (images) that have a high similarity to the "weighted query-corresponding feature vector 131" calculated by the weighted query-corresponding feature vector generation unit 122, and outputs the selected content (images) to the output unit 112 as search results 133. The search results 133 are displayed in the search result display area 13 of the data search UI shown in FIG. 1 .
[0253] The feature vector similarity analysis unit 123 may perform processing such as limiting the highly similar content (image) to be output to the output unit 112 as the search result 133 to content (image) having a similarity equal to or greater than a preset threshold value. Also, the feature vector similarity analysis unit 123 may perform processing such as sorting the highly similar content (image) to be output to the output unit 112 as the search result 133 in descending order of similarity, and displaying the content (image) in descending order of similarity.
[0254] Furthermore, as shown in FIG. 18, a content encoder 124 may be provided inside the data processing unit 102 of the information processing device 100, and the content such as an image input via the communication unit 103 may be analyzed (encoded) to calculate a content-corresponding feature vector.
[0255] The content encoder 124 uses a deep learning model (Deep Neural Network) that performs image analysis and calculates feature vectors, similar to the content encoder 211 described above with reference to Fig. 17. Specifically, similar to the above-mentioned Transformer, a Vision Transformer (ViT), a Swin Transformer, or a ResNet that performs image analysis using an attention mechanism can be used.
[0256] 5. Processing Sequence Executed by the Information Processing Device of the Present Disclosure Next, a processing sequence executed by the information processing device of the present disclosure will be described.
[0257] The sequence of processing executed by the information processing device of the present disclosure will be described with reference to the flowcharts shown in FIGS.
[0258] 19 and 20 are executed in accordance with a program stored in a storage unit of the information processing device of the present disclosure under the control of a data processing unit including a CPU having a program execution function, etc. The processing of each step of the flow shown in FIGS.
[0259] (Step S101) First, in step S101, the data processing unit 102 of the information processing device 100 of the present disclosure starts a search process execution screen (UI) and displays it on the display unit.
[0260] For example, the data search UI shown in FIG. 1 is displayed on the display unit of the information processing device 100 .
[0261] (Step S102) Next, in step S102, the data processing unit 102 of the information processing device 100 of the present disclosure inputs a search query (search statement).
[0262] For example, a search query (search statement) entered by a user in the query input area 11 in the data search UI shown in FIG.
[0263] (Step S103) Next, in step S103, the data processing unit 102 of the information processing apparatus 100 of the present disclosure determines whether or not a weight is set for the input search query (search statement).
[0264] That is, it is determined whether or not the search query (search sentence) is a query in which weights are set for each constituent word. Note that a query in which weights are set for each constituent word is a query in which any of the following weight settings have been made, as described with reference to FIG. 2 and subsequent figures: (a) Weight setting by inputting a numerical value (b) Weight setting by inputting brackets (c) Weight setting using a scroll bar (d) Weight setting using a color display (heat map)
[0265] If it is determined that the input query (search statement) is a query without weight setting, the process proceeds to step S111. On the other hand, if it is determined that the input query (search statement) is a query with weight setting, the process proceeds to step S121.
[0266] (Step S111) If it is determined in step S103 that the input query (search statement) is a query without weight setting, the following process is executed in step S111.
[0267] In step S111, the data processing unit 102 executes an encoding process on the input search query (search statement) to calculate a query-related feature vector.
[0268] This process is executed by the weighted query-corresponding feature vector generation unit 122 of the data processing unit 102 shown in Fig. 14. Note that this process of calculating the query-corresponding feature vector can be performed using, for example, the unidirectional attention-type Transformer previously described with reference to Fig. 15(a). That is, this is the process of calculating the "query-corresponding feature vector (Vec)" shown in Fig. 15(a), and is a vector generated by the text encoder 121 of the data processing unit 102 shown in Fig. 14.
[0269] The weighted query-corresponding feature vector generation unit 122 of the data processing unit 102 shown in Fig. 14 outputs the "query-corresponding feature vector (Vec)" input from the text encoder 121 as is as the "weighted query-corresponding feature vector 131" shown in Fig. 14. However, the "weighted query-corresponding feature vector 131" in this case is a feature vector in which the weights of all words are set to 1.0, and is a vector equivalent to the unweighted "query-corresponding feature vector (Vec)."
[0270] The weighted query-corresponding feature vector generation unit 122 outputs the "query-corresponding feature vector (Vec)" input from the text encoder 121 as is, thereby making it possible to omit the calculation processes of steps S01 to S02 described above with reference to FIG. 16 .
[0271] However, even in the case of a query (search sentence) for which no weight is set, the calculation process of steps S01 to S02 described above with reference to Fig. 16 may be executed. That is, the weight s x are all set to 1.0, and the weight-set query-corresponding feature vector V′ is calculated according to the following formula (Formula 2) described above. K " 131 is calculated. V' K =V 0 +s 1 D 1 +...+s K D K ...(Formula 2)
[0272] If the user does not input the weight (scale) for each word (x), all the weights (s x )=1.0. In this case, the weighted query-corresponding feature vector V′ in (Equation 2) above is K The calculation formula is as follows (Formula 3): V' K =V 0 +D 1 +...+D K ...(Formula 3)
[0273] In the above formula (3), D 1 =V 1 -V 0 D2 =V 2 -V 1 : D K =V K -V K-1 Therefore, the above equation (3) becomes: V' K =V 0 +D 1 +...+D K =V 0 + (V 1 -V 0 ) + (V 2 -V 1 ) + ... (V K -V K-1 ) = V K The "weighted query-corresponding feature vector V'" calculated by the above (Equation 3) is K " is a vector that matches the "query-corresponding feature vector (Vec)" shown in (a) of FIG. 15(a).
[0274] In this way, when a query (search statement) without weights is input, the weighted query-responsive feature vector generation unit 122 performs the following processing: (a) outputs the "query-responsive feature vector (Vec)" input from the text encoder 121 as is, or (b) outputs the weight s for each word (x) input by the user. x are all set to 1.0, and the weight-set query-corresponding feature vector V′ is calculated according to the following formula (Formula 2) described above. K " 131 is calculated. V' K =V 0 +s 1 D 1 +...+s K D K ... (Equation 2) Any configuration may be used as long as it executes either the above process (a) or (b).
[0275] In step S111, when the calculation process of the query-related feature vector corresponding to the query (search statement) without weight setting is completed, the process proceeds to step S112.
[0276] (Step S112) Next, in step S112, the data processing unit 102 of the information processing device 100 of the present disclosure acquires or calculates a feature vector corresponding to the search target content (image, etc.).
[0277] 14 and 17, content IDs and feature vectors are stored together with content such as images in the databases 200-1 to 200-n that store content such as images accessible by the information processing device 100 via the communication unit 103. In step S112, the data processing unit 102 of the information processing device 100 of the present disclosure acquires the stored feature vectors associated with the images from the databases 200-1 to 200-n that store the content.
[0278] In addition, if a feature vector associated with an image is not stored in the database 200-1 to 200-n that stores the content, the content encoder 124 in the data processing unit 102 of the information processing device 100 analyzes the image acquired from the database 200-1 to 200-n and performs the feature vector calculation process, as previously described with reference to Figure 18.
[0279] (Step S113) Next, in step S113, the data processing unit 102 of the information processing device 100 of the present disclosure analyzes the similarity between the query-corresponding feature vector calculated in step S111 and the content (image, etc.)-corresponding feature vector acquired or calculated in step S112.
[0280] As described above, the similarity analysis process between these feature vectors utilizes an index value of similarity between vectors, such as cosine similarity or inter-vector distance (Euclidean distance).
[0281] (Step S131) Next, in step S131, the data processing unit 102 of the information processing device 100 of the present disclosure selects content (image) having a content-corresponding feature vector whose similarity to the query-corresponding feature vector is equal to or greater than a specified threshold, based on the result of the similarity analysis of the feature vectors in step S113.
[0282] (Step S132) Next, in step S132, the data processing unit 102 of the information processing device 100 of the present disclosure sorts the contents (images) having content-corresponding feature vectors whose similarity to the query-corresponding feature vector selected in step S131 is equal to or greater than a specified threshold value as search results, and outputs the contents (images) to the display unit of the information processing device 100 in descending order of similarity.
[0283] That is, these images are displayed as search results in the search result display area 13 of the data search UI displayed on the display unit of the information processing apparatus described above with reference to FIG.
[0284] Next, the process following step S121 that is executed when the determination in step S103 shown in FIG. 19 is Yes, that is, when it is determined that the input query (search statement) is a query with weight setting, will be described.
[0285] (Step S121) If it is determined in step S103 that the input query (search statement) is a query with weight setting, the following process is executed in step S121.
[0286] First, in step S121, an encoding process is executed on the input search query (search statement) to calculate a token correspondence vector. This process is the process of the text encoder 121 described above with reference to FIG. 16. As shown in the lower left corner of FIG. 16, the text encoder 121 calculates the token correspondence vector (V 0 ~V N ) is calculated.
[0287] (Step S122) Next, in step S122, the data processing unit 102 performs an arithmetic process to apply the token correspondence vector calculated by the text encoder 121 in step S121 and the weight set by the user at the word level to calculate a weighted query correspondence feature vector.
[0288] This process is a calculation process executed by the weighted query-responsive feature vector generation unit 122 described above with reference to FIG. 16, and corresponds to steps S01 to S02 shown in FIG.
[0289] The weighted query-corresponding feature vector generation unit 122 first calculates the difference between adjacent tokens in the “token correspondence vector” generated by the text encoder 121 as a token correspondence vector difference (D X ) is calculated as the token corresponding vector difference (D X ) into the adjacent token correspondence vector (V X ~V X-1 The difference (Diff) between Dx and V is calculated according to the following formula (Formula 1): Dx = V X -V X-1 ...(Formula 1)
[0290] Next, the weight (scale) of each word (x) input by the user is calculated as s x Then, according to the following formula (Formula 2), a "weighted query-corresponding feature vector V' K " is calculated. K =V 0 +s 1 D 1 +...+s K D K ...(Formula 2)
[0291] The weighted query-responsive feature vector generation unit 122 performs these calculations to generate a weighted query-responsive feature vector V' K " is calculated.
[0292] (Step S123) Next, in step S123, the data processing unit 102 of the information processing device 100 of the present disclosure acquires or calculates a feature vector corresponding to the search target content (image, etc.).
[0293] This process is the same as the process of step S112 described above. That is, the data processing unit 102 of the information processing device 100 disclosed herein acquires a stored feature vector associated with an image from the databases 200-1 to 200-n that store content. Alternatively, if the databases 200-1 to 200-n do not store a stored feature vector associated with an image, the content encoder 124 in the data processing unit 102 of the information processing device 100 analyzes the image acquired from the databases 200-1 to 200-n and calculates the feature vector, as described above with reference to FIG. 18 .
[0294] (Step S124) Next, in step S124, the data processing unit 102 of the information processing device 100 of the present disclosure analyzes the similarity between the weighted query-corresponding feature vector calculated in step S122 and the content (image, etc.)-corresponding feature vector acquired or calculated in step S123.
[0295] As described above, the similarity analysis process between these feature vectors utilizes an index value of similarity between vectors, such as cosine similarity or inter-vector distance (Euclidean distance).
[0296] (Steps S131 to S132) The processes in steps S131 to S132 are the same as those described above. That is, in step S131, based on the result of the similarity analysis of the feature vectors in step S124, content (images) having a content-corresponding feature vector whose similarity to the weighted query-corresponding feature vector is equal to or greater than a specified threshold value are selected.
[0297] Next, in step S132, the search results are sorted by decreasing similarity and output to the display unit of the information processing device 100, and the content (images) having content-corresponding feature vectors whose similarity to the weighted query-corresponding feature vector selected in step S131 is equal to or greater than a specified threshold value.
[0298] That is, these images are displayed as search results in the search result display area 13 of the data search UI displayed on the display unit of the information processing apparatus described above with reference to FIG.
[0299] By performing such processing, it is possible to obtain search results that reflect the weights of individual words set by the user, and to efficiently obtain the search results that the user intended.
[0300] In the above-described embodiment, an example has been described in which image data is used as the content to be searched, but the content or data to be searched is not limited to images and may be data of other categories. In other words, the category of data to be searched by the information processing device of the present disclosure is not limited to images. For example, it is possible to perform a data search using data of various categories, such as documents, music, architecture, nature, and living things, in addition to images.
[0301] 6. Application Examples of Data Search Processing Using the Information Processing Device of the Present Disclosure Next, application examples of data search processing using the information processing device of the present disclosure will be described.
[0302] As application examples of data search processing using the information processing device of the present disclosure, the following three processing examples will be described: (1) Processing example of extracting a specific scene intended by a user from movie content; (2) Processing example of extracting a specific scene intended by a user from sports video content; and (3) Processing example of extracting motion data intended by a user from motion data.
[0303] (6-(1) Example of Processing for Extracting a Specific Scene Intended by a User from Movie Content) First, an example of processing for extracting a specific scene intended by a user from movie content will be described.
[0304] Movies are often distributed around the world, but there are various regulations depending on the country or region. For example, there are cases where violent scenes are requested to be deleted. However, even when it comes to violent scenes, the standards vary from country to country and region, so it is not possible to apply a uniform standard.
[0305] In such cases, by performing the process of the present disclosure described above, i.e., scene extraction in which weights are assigned to the words that make up the query (search sentence), it is possible to extract scenes that meet the criteria of each region.
[0306] For example, as shown in Fig. 21, the following query (search statement) is input: Query = extreme or violent scene. The above query is then entered, and a search is performed by setting various weights for the words "extreme" and "violent" contained in the query. By performing this type of processing, it becomes possible to extract scenes that meet the standards of various regions.
[0307] For example, when extracting scenes in accordance with the criteria of area A, as shown in Figure 21, by setting the weights as follows: "Extreme" weight = 1.5 times "Violent" weight = 1.8 times, scene extraction can be performed in accordance with the criteria of area A.
[0308] Furthermore, when extracting scenes in accordance with the standards of area B, which differ from area A, the weighting settings are changed. For example, weighting for "extreme" = 1.0 times, weighting for "violent" = 1.2 times. By setting the weighting in this way, it is possible to extract scenes in accordance with the standards of area B.
[0309] The processing for performing scene extraction processing tailored to the actual conditions of such a region will be described with reference to Fig. 22. The data processing unit 102 shown in Fig. 22 has the configuration of the data processing unit 102 of the information processing device 100 of the present disclosure previously described with reference to Fig. 14 and Fig. 18. Note that the content encoder 124 is indicated by a dotted line because it may or may not be used.
[0310] Various movie contents are stored in the movie database 210. If the movie database 210 already stores scene-corresponding feature vectors associated with each scene of the movie contents, the content encoder 124 is not necessary, and the scene-corresponding feature vectors 132-1 to 132-n are acquired from the movie database 210.
[0311] On the other hand, if the movie database 210 does not store scene-corresponding feature vectors associated with each scene of the movie content, the content encoder 124 analyzes each scene of the movie content and calculates the scene-corresponding feature vectors 132-1 to 132-n.
[0312] The text encoder 121 and the weighted query-corresponding feature vector generation unit 122 calculate a "weighted query-corresponding feature vector 131" using the query (search statement) input via the input unit 111, i.e., Query = extreme or violent scene, and the word-by-word weight setting information described with reference to Figure 21.
[0313] The vector similarity analysis unit 123 analyzes the similarity between the "weighted query-corresponding feature vector 131" generated by the weighted query-corresponding feature vector generation unit 122 and the scene-corresponding feature vectors 132-1 to 132-n, and outputs scenes with high similarity as search results to the output unit 112.
[0314] The user checks the output results, and if the search results they intended are not obtained, they change the weight settings of the words that make up the query and perform the search again. By repeating this process, the user can obtain the search results they intended.
[0315] The flowchart shown in Fig. 23 is a flowchart for explaining the sequence of the above-mentioned processing. The processing of each step in the flowchart shown in Fig. 23 will be explained in order.
[0316] (Step S201) First, the data processing unit 102 of the information processing device 100 of the present disclosure inputs the following query (search sentence) via the input unit 111: Query=Extreme or Violent Scene The above query and weight information for each word.
[0317] For example, the query and weight information are input using the UI described with reference to FIG.
[0318] (Step S202) Next, the data processing unit 102 of the information processing device 100 of the present disclosure generates a weighted query-corresponding feature vector based on the input query and weight information for each word, and executes a data search process to extract movie scenes having scene-corresponding feature vectors that have a high similarity to the generated weighted query-corresponding feature vector.
[0319] (Step S203) Next, the data processing unit 102 of the information processing device 100 of the present disclosure outputs, as search results, to the output unit 112, movie scenes having scene-corresponding feature vectors that have high similarity to the weighted query-corresponding feature vector acquired in the search process of step S202.
[0320] These search results are displayed in the search result display area 13 of the data search UI displayed on the display unit of the information processing apparatus described above with reference to FIG.
[0321] (Step S204) Step S204 is a process in which the user checks the search results. If it is determined that search results that match the user's intention have been obtained, the search process ends.
[0322] On the other hand, if it is determined that no search results matching the user's intention have been obtained, the process proceeds to step S205.
[0323] (Step S205) If it is determined in step S204 that no search results matching the user's intention have been obtained, the process of step S205 is executed.
[0324] The user can input the following query as is: Query = extreme or violent scenes, or change the weight information for each word.
[0325] Thereafter, in step S204, the processes from S202 onward are repeated. By repeating the weight change process and search process in this manner, if the determination in step S204 is Yes, that is, if it is determined that the search results intended by the user have been obtained, the process ends.
[0326] In this way, by applying the processing of the present disclosure, it is possible to adjust the weights of each constituent word of a query and efficiently obtain the search results that the user intends.
[0327] (6-(2) Example of Processing for Extracting a Specific Scene Intended by a User from Sports Video Content) Next, an example of processing for extracting a specific scene intended by a user from sports video content will be described.
[0328] For example, when creating a highlight video by extracting only notable scenes from video content of a soccer match, it is necessary to process the following scenes: (a) The scene where a goal is scored, (b) The scene where the player strikes a decisive pose after scoring a goal, and (c) The scene where the goalkeeper makes a spectacular save.
[0329] For example, when extracting such scenes from video content of a soccer match, applying the processing of the present disclosure makes it possible to efficiently extract scenes that match the user's intentions.
[0330] For example, as shown in FIG. 24, the following query (search statement) is input: Query = Scene of scoring a goal and striking a decisive pose. The above query is input, and a search is performed by setting various weights to the words (word strings) contained in this query, namely "scoring a goal" and "striking a decisive pose."
[0331] For example, as shown in FIG. 24, the weight of "scoring a goal" = 1.5 times, and the weight of "finishing a pose" = -0.2 times. By setting the weights in this way, scenes can be extracted that focus on scenes where goals are scored.
[0332] By performing such processing, it becomes possible to extract scenes that match the user's intention. Note that the information processing device 100 of the present disclosure is capable of setting weights not only for a single word but also for a word string consisting of multiple words.
[0333] A processing configuration for performing such scene extraction processing will be described with reference to Fig. 25. The data processing unit 102 shown in Fig. 25 has the configuration of the data processing unit 102 of the information processing device 100 of the present disclosure previously described with reference to Fig. 14 and Fig. 18. Note that the content encoder 124 is indicated by a dotted line because it may or may not be used.
[0334] Various sports video contents are stored in the sports video database 220. If the sports video database 220 already stores scene-corresponding feature vectors associated with each video scene, the content encoder 124 is not necessary, and the scene-corresponding feature vectors 132-1 to 132-n are acquired from the sports video database 220.
[0335] On the other hand, if the sports video database 220 does not store scene-corresponding feature vectors associated with each video scene, the content encoder 124 analyzes each scene of the sports video content and calculates scene-corresponding feature vectors 132-1 to 132-n.
[0336] The text encoder 121 and the weighted query-corresponding feature vector generation unit 122 calculate a "weighted query-corresponding feature vector 131" by using the query (search statement) input via the input unit 111, i.e., Query == Scene of scoring a goal and striking a pose, and the weight setting information for each word (word string) described with reference to Figure 24.
[0337] The vector similarity analysis unit 123 analyzes the similarity between the "weighted query-corresponding feature vector 131" generated by the weighted query-corresponding feature vector generation unit 122 and the scene-corresponding feature vectors 132-1 to 132-n, and outputs scenes with high similarity as search results to the output unit 112.
[0338] The user checks the output results, and if the search results they intended are not obtained, they change the weight settings of the words that make up the query and perform the search again. By repeating this process, the user can obtain the search results they intended.
[0339] (6-(3) Example of Processing for Extracting User-Intended Motion Data from Motion Data) Next, an example of processing for extracting user-intended motion data from motion data will be described.
[0340] For example, when performing processing to make a character in game content or the like perform a specific movement, it is possible to obtain the motion data intended by the user from a motion database that stores various types of motion data, and then use the obtained motion data to control the character's movement.
[0341] However, a large number of different motion data are stored in the motion database, and it is difficult to search for motion data that matches the user's intended movement.
[0342] 26 shows an example of motion data stored in the motion database 230. The examples shown in the figure are examples of "running" motion data and "running and jumping" motion data. In addition to these, a large amount of motion data is stored in the motion database 230. By applying the processing of the present disclosure, it is possible to efficiently extract the motion data intended by the user.
[0343] 27 shows two examples of queries (search statements) used for motion data searches. The query (search statement) in (a) is: Query = Motion data of running as if tired, and is an example of a search performed by setting a weight on the word "as if tired" included in the query.
[0344] The query (search statement) in (a) is: Query = Motion data of a quick running movement, and is an example of a search performed by setting a weight on the word "quickly" included in the query.
[0345] By adjusting the weight of each query, it is possible to select motion data with different degrees of fatigue in case (a), and motion data with different degrees of quickness in case (b). By performing such processing, it is possible to efficiently extract motion data that matches the user's intentions.
[0346] The processing configuration for performing such motion data extraction processing will be described with reference to Fig. 28. The data processing unit 102 shown in Fig. 28 is the configuration of the data processing unit 102 of the information processing device 100 of the present disclosure previously described with reference to Fig. 14 and Fig. 18. Note that the content encoder 124 is indicated by a dotted line because it may or may not be used.
[0347] Motion data for various movements is stored in the motion data database 230. If motion data corresponding feature vectors associated with each piece of motion data are already stored in the motion data database 230, the content encoder 124 is not necessary, and the motion data corresponding feature vectors 132-1 to 132-n are acquired from the motion data database 230.
[0348] On the other hand, if the motion data database 230 does not store motion data-associated feature vectors associated with the motion data, the content encoder 124 analyzes the motion data and calculates motion data-associated feature vectors 132-1 to 132-n.
[0349] The text encoder 121 and the weighted query-corresponding feature vector generation unit 122 calculate a "weighted query-corresponding feature vector 131" using a query (search statement) input via the input unit 111, for example, Query == Motion data of tired-looking running, and the weight setting information for each word (word string) described with reference to Figure 27.
[0350] The vector similarity analysis unit 123 analyzes the similarity between the "weighted query-corresponding feature vector 131" generated by the weighted query-corresponding feature vector generation unit 122 and the motion data-corresponding feature vectors 132-1 to 132-n, and outputs the motion data with high similarity as the search result to the output unit 112.
[0351] The user checks the output results, and if the search results they intended are not obtained, they change the weight settings of the words that make up the query and perform the search again. By repeating this process, the user can obtain the motion data they intended as search results.
[0352] 7. Hardware Configuration Example of Information Processing Device of the Present Disclosure Next, a hardware configuration example of an information processing device of the present disclosure will be described with reference to FIG.
[0353] 29 corresponds to hardware of a PC, a smartphone, a tablet terminal, etc., which are examples of information processing devices of the present disclosure. Each component constituting the hardware of the information processing device shown in FIG. 29 will be described below.
[0354] The CPU (Central Processing Unit) 301 functions as a control unit or data processing unit that executes various processes according to programs stored in the ROM (Read Only Memory) 302 or the storage unit 308. For example, it executes processes according to the sequences described in the above-mentioned embodiments. The RAM (Random Access Memory) 303 stores programs and data executed by the CPU 301. The CPU 301, ROM 302, and RAM 303 are interconnected by a bus 304.
[0355] The CPU 301 is connected to an input / output interface 305 via a bus 304, and the input / output interface 305 is connected to an input unit 306 including various switches, buttons, a touch panel, sensors, a camera, a microphone, etc., and an output unit 307 that outputs data to a display unit, a speaker, etc. The CPU 301 executes various processes in response to commands input from the input unit 306, and outputs the processed results to the output unit 307, for example.
[0356] A storage unit 308 connected to the input / output interface 305 stores various data and programs executed by the CPU 301. A communication unit 309 functions as a transmitter / receiver for data communication via Wi-Fi communication, Bluetooth (registered trademark) communication, or other networks such as the Internet or a local area network, and communicates with external devices.
[0357] A drive 310 connected to the input / output interface 305 drives a removable medium 311 such as a semiconductor memory such as a memory card, and executes recording or reading of data.
[0358] [8. Summary of the Configuration of the Present Disclosure] The embodiments of the present disclosure have been described above in detail with reference to specific examples. However, it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the gist of the present disclosure. In other words, the present invention has been disclosed in the form of examples and should not be interpreted as being limited. To determine the gist of the present disclosure, the claims should be taken into consideration.
[0359] The technology disclosed in this specification can be configured as follows: (1) An information processing device that acquires a query made up of text to be applied to data search processing, and has a data processing unit that executes data search processing based on the query, wherein the data processing unit acquires, together with the query, weight information for each word or word string that constitutes the query, and executes search processing for data in a category different from text by using the query and the weight information for each word or word string.
[0360] (2) The information processing device according to (1), wherein the data in a category different from text is either image data or music data.
[0361] (3) The information processing device according to (1) or (2), wherein the weight information for each word or word string constituting the query is a positive weight or a negative weight, and the data processing unit executes a search process with an increased influence in the search process for words or word strings to which a positive weight is set, and executes a search process with a decreased influence in the search process for words or word strings to which a negative weight is set.
[0362] (4) The information processing device is an information processing device described in any one of (1) to (3), wherein the information processing device has a display unit that displays a user interface for data search, and the data processing unit displays a user interface for data search on the display unit that has a query input area that allows weight information for each word or word string to be input.
[0363] (5) The information processing device according to (4), wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input as a numerical value.
[0364] (6) The information processing device according to (4) or (5), wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input using brackets.
[0365] (7) The information processing device according to any one of (4) to (6), wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input by operating a scroll bar.
[0366] (8) The information processing device according to any one of (4) to (7), wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input using a color display (heat map).
[0367] (9) The information processing device described in any one of (1) to (8), wherein the data processing unit uses the query and weight information for each word or word string to generate a weighted query-corresponding feature vector, and executes a content search process to select and acquire content having a content-corresponding feature vector that is highly similar to the generated weighted query-corresponding feature vector.
[0368] (10) The information processing device according to (9), wherein the data processing unit includes: an encoder that generates a token correspondence vector corresponding to each token, which is a processing unit including constituent words of the query; and a weighted query correspondence feature vector generation unit that performs a calculation process using the token correspondence vectors generated by the encoder to generate a weighted query correspondence feature vector.
[0369] (11) The information processing device according to (10), wherein the encoder is an encoder that uses a deep learning model that uses an attention mechanism, and has a unidirectional attention mechanism that analyzes, for each word that constitutes the query, only the relevance of that word with words that follow that word.
[0370] (12) The information processing device according to (10) or (11), wherein the weighted query-corresponding feature vector generation unit calculates a token correspondence vector difference, which is a difference between adjacent tokens of the token correspondence vector generated by the encoder, and performs an arithmetic process on the calculated token correspondence vector difference and weight information for each word or word string constituting the query, to generate a weighted query-corresponding feature vector.
[0371] (13) The weighted query-corresponding feature vector generation unit generates a token correspondence vector difference (D X ) and weight information (s x ) and V' K =V 0 +s 1 D 1 +...+s K D KHowever, V' K is the weighted query-aware feature vector, V 0 is the leading token correspondence vector output by the encoder, s n is the weight assigned to the n-th word in the query, D n is the difference between the token correspondence vector of the nth word and the token correspondence vector of the (n-1)th word among the constituent words of the query. By performing the calculation process according to the above formula, a weighted query correspondence feature vector V' is obtained. K The information processing device according to any one of (10) to (12) above,
[0372] (14) The information processing device according to any one of (9) to (13), wherein the data processing unit has a feature vector similarity analysis unit for selecting content having a content-corresponding feature vector that is highly similar to the weight-setting query-corresponding feature vector.
[0373] (15) The information processing device according to (14), wherein the feature vector similarity analysis unit calculates a cosine similarity or a Euclidean distance, which is a vector similarity determination index value, in a similarity analysis process between the weighted query-corresponding feature vector and the content-corresponding feature vector, to perform vector similarity analysis.
[0374] (16) The information processing device according to any one of (1) to (15), wherein the data processing unit executes a search process using stored data in a database as search target data.
[0375] (17) The information processing device described in (16), wherein the database stores content to be searched and content-corresponding feature vectors generated by feature analysis of the content, and the data processing unit executes a content search process to select and acquire from the database content having a content-corresponding feature vector that is highly similar to the weighted query-corresponding feature vector generated using the query and weight information on a word or word string basis.
[0376] (18) The information processing device according to (16) or (17), wherein the database stores content to be searched, and the data processing unit performs a content search process to analyze the content stored in the database to calculate a content-corresponding feature vector, select from the calculated content-corresponding feature vectors a content-corresponding feature vector that has a high similarity to a weighted query-corresponding feature vector generated using the query and weight information on a word or word string basis, and select and acquire from the database as a search result the content for which the selected content-corresponding feature vector has been generated.
[0377] (19) An information processing method that acquires a query composed of text to be applied to a data search process, and weight information for each word or word string that constitutes the query, and performs a search process for data in a category different from the text by using the query and the weight information for each word or word string.
[0378] (20) A program for executing information processing in an information processing device, the program causing a data processing unit of the information processing device to acquire a query composed of text to be applied to data search processing and weight information for each word or word string that constitutes the query, and executing search processing for data in a category different from the text using the query and the weight information for each word or word string.
[0379] The series of processes described in this specification can be executed by hardware, software, or a combination of both. When executing processes by software, a program recording the processing sequence can be installed and executed in the memory of a computer incorporated in dedicated hardware, or the program can be installed and executed on a general-purpose computer capable of executing various processes. For example, the program can be pre-recorded on a recording medium. In addition to installing the program on a computer from a recording medium, the program can also be received via a network such as a LAN (Local Area Network) or the Internet and installed on a recording medium such as an internal hard disk.
[0380] The various processes described in this specification may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capabilities of the devices executing the processes or as needed. Furthermore, in this specification, a system refers to a logical collective configuration of multiple devices, and is not limited to devices that are all located in the same housing.
[0381] As described above, according to the configuration of one embodiment of the present disclosure, an apparatus and method are realized that perform search processing for data in a category different from text using a query and weight information for each word. Specifically, for example, the apparatus includes a data processing unit that inputs a query composed of text to be applied to the data search processing and performs data search processing based on the query. The data processing unit inputs weight information for each word that constitutes the query along with the query, and performs search processing for data in a category different from text using the query and the weight information for each word. For words assigned with positive weights, search processing is performed with an increased influence, and for words assigned with negative weights, search processing is performed with a decreased influence. This configuration realizes an apparatus and method that perform search processing for data in a category different from text using a query and weight information for each word.
[0382] 11 Query input area 12 Search (start search) icon 13 Search result (Result) display area 100 Information processing device 101 UI (user interface) 102 Data processing unit 103 Communication unit 111 Input unit 112 Output unit 121 Text encoder 122 Weighted query-corresponding feature vector generation unit 123 Similarity analysis unit 124 Content encoder 131 Weighted query-corresponding feature vector 132 Content-corresponding feature vector 133 Search results 200, 210, 220, 230 Database 211 Content encoder 301 CPU 302 ROM 303 RAM 304 Bus 305 Input / output interface 306 Input unit 307 Output unit 308 Storage unit 309 Communication unit 310 Drive 311 Removable media
Claims
1. An information processing device having a data processing unit that acquires a query composed of text to be applied to data search processing, and executes data search processing based on the query, wherein the data processing unit acquires weight information for each word or word string that constitutes the query along with the query, and executes search processing for data in a category different from the text using the query and the weight information for each word or word string.
2. The information processing device according to claim 1, wherein the data in a category different from text is either image data or music data.
3. The information processing device of claim 1, wherein the weight information for each word or word string constituting the query is a positive weight or a negative weight, and the data processing unit executes a search process with an increased influence in the search process for words or word strings to which a positive weight is set, and executes a search process with a decreased influence in the search process for words or word strings to which a negative weight is set.
4. The information processing device according to claim 1, wherein the information processing device has a display unit that displays a user interface for data search, and the data processing unit displays a user interface for data search on the display unit that has a query input area in which weight information for each word or word string can be input.
5. The information processing device according to claim 4, wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input as a numerical value.
6. The information processing device according to claim 4, wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input using brackets.
7. The information processing device according to claim 4, wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input by operating a scroll bar.
8. The information processing device according to claim 4, wherein the data processing unit displays on the display unit a user interface for data search having a query input area in which weight information for each word or word string can be input using a color display (heat map).
9. The information processing device according to claim 1, wherein the data processing unit generates a weighted query-corresponding feature vector using the query and weight information for each word or word string, and executes a content search process to select and acquire content having a content-corresponding feature vector that is highly similar to the generated weighted query-corresponding feature vector.
10. The information processing device according to claim 9, wherein the data processing unit has an encoder that generates a token correspondence vector corresponding to each token, which is a processing unit including constituent words of the query, and a weighted query correspondence feature vector generation unit that performs calculation processing using the token correspondence vectors generated by the encoder to generate a weighted query correspondence feature vector.
11. The information processing device described in claim 10, wherein the encoder is an encoder that uses a deep learning model that utilizes an attention mechanism, and has a unidirectional attention mechanism that analyzes only the relevance of each word that constitutes the query with words that follow that word.
12. The information processing device according to claim 10, wherein the weighted query-corresponding feature vector generation unit calculates a token correspondence vector difference, which is a difference between adjacent tokens in the token correspondence vector generated by the encoder, and performs an arithmetic operation on the calculated token correspondence vector difference and weight information for each word or word string constituting the query to generate a weighted query-corresponding feature vector.
13. The weight-set query-corresponding feature vector generation unit generates a token correspondence vector difference (D X ) and weight information (s x ) and V' K =V 0 +s 1 D 1 +...+s K D K However, V' K is the weighted query-aware feature vector, V 0 is the leading token correspondence vector output by the encoder, s n is the weight assigned to the n-th word in the query, D n is the difference between the token correspondence vector of the nth word and the token correspondence vector of the (n-1)th word among the constituent words of the query. By performing the calculation process according to the above formula, a weighted query correspondence feature vector V' is obtained. K The information processing apparatus according to claim 10 , wherein the information processing apparatus calculates:
14. The information processing device according to claim 9, wherein the data processing unit has a feature vector similarity analysis unit for selecting content having a content-corresponding feature vector that is highly similar to the weight-setting query-corresponding feature vector.
15. The information processing device according to claim 14, wherein the feature vector similarity analysis unit calculates cosine similarity or Euclidean distance, which is a vector similarity determination index value, in the similarity analysis process between the weighted query-corresponding feature vector and the content-corresponding feature vector, to perform vector similarity analysis.
16. The information processing device according to claim 1, wherein the data processing unit executes a search process using data stored in a database as search target data.
17. An information processing device as described in claim 16, wherein the database stores content to be searched and content-corresponding feature vectors generated by feature analysis of the content, and the data processing unit executes a content search process to select and acquire from the database content having a content-corresponding feature vector that is highly similar to a weighted query-corresponding feature vector generated using the query and weight information on a word or word string basis.
18. An information processing device according to claim 16, wherein the database stores content to be searched, and the data processing unit executes a content search process in which it analyzes the content stored in the database to calculate a content-corresponding feature vector, selects from the calculated content-corresponding feature vectors a content-corresponding feature vector that has a high similarity to a weighted query-corresponding feature vector generated using the query and weight information on a word or word string basis, and selects and acquires from the database as a search result the content for which the selected content-corresponding feature vector has been generated.
19. An information processing method that acquires a query composed of text to be applied to data search processing and weight information for each word or word string that constitutes the query, and performs search processing for data in a category different from the text by using the query and the weight information for each word or word string.
20. A program for executing information processing in an information processing device, which causes a data processing unit of the information processing device to acquire a query composed of text to be applied to data search processing and weight information for each word or word string that constitutes the query, and executes search processing for data in a category different from the text using the query and the weight information for each word or word string.
Citation Information
Patent Citations
Retrieval system, terminal apparatus, information processing apparatus, retrieval method and program
JP2018194903A
Automatically curated image searching
US20190163768A1