Information processing system, information processing apparatus, information processing method, and program
Patent Information
- Application Number
- JP2024187517
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2024-10-24
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2044-10-24
AI Technical Summary
It is difficult for the prior art to effectively provide information in line with user preferences to recommend products or services.
An information processing system is designed, including electronic devices and information processing equipment. The system extracts image-related feature and tag information by searching and using images that use user preferences to select, combining image databases and text information processing, and provides personalized information recommendations.
It realizes the provision of personalized information based on user preferences, improving the accuracy and user experience of information recommendations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing system, an information processing device, an information processing method, and a program for executing image-related processing. [Background technology]
[0002] Conventionally, there are techniques for executing various processes using various information related to images. For example, a technique for presenting a plurality of images for introducing various products to a user on a website has been proposed (for example, see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-082053 Summary of the Invention [Problem to be solved by the invention]
[0004] For example, in order to understand a user's abstract ideal image and preferences for the design taste of a product or service and recommend such a product or service to the user, it is important to provide the user with information that matches the user's preferences.
[0005] An object of the present invention is to appropriately provide information according to a user's preferences. [Means for solving the problem]
[0006] One aspect of the present invention is an information processing system including an electronic device used by a user and an information processing device capable of providing the electronic device with recommended images searched using selected images selected according to the user's preferences, the information processing system including a database that stores images and features of the images in association with each of a plurality of images, an acquisition unit that acquires text information indicating a user's desire for the recommended images, a feature generation unit that generates a first feature related to the user's preferences based on the text information and generates a second feature that is a feature of the selected image related to what is contained in the selected image based on the selected image, and a selection unit that selects a new image according to the user's preferences from among the images stored in the database based on a comparison result between the feature information indicating the user's preferences generated based on the first feature and the second feature and the features of the image, an information processing device that constitutes the information processing system, an information processing method including each of those processes, and a program that causes a computer to execute each of those processes. Effect of the Invention
[0007] According to the present invention, it is possible to appropriately provide information according to the preferences of a user. [Brief description of the drawings]
[0008] [Figure 1] 2 is a block diagram showing an example of a functional configuration of an information processing device; [Diagram 2] 2 is a block diagram showing an example of a functional configuration of an information processing device; [Diagram 3] FIG. 13 is a diagram showing the contents of a tag list stored in a tag DB. [Figure 4] FIG. 2 is a diagram showing the stored contents of image information and tag information in an image DB. [Diagram 5] FIG. 13 is a diagram showing the flow of a setting method for setting a plurality of tags. [Figure 6] FIG. 11 is a diagram showing a flow of an extraction process for extracting a feature amount for each tag. [Figure 7] 13 is a flowchart illustrating an example of a tag setting process. [Figure 8] 13 is a flowchart illustrating an example of a feature extraction process. [Figure 9] 2 is a block diagram showing an example of a functional configuration of an information processing device; [Figure 10] FIG. 13 is a diagram showing the contents of weight data stored in a weight DB. [Figure 11] FIG. 11 is a diagram illustrating a flow of a weight calculation process. [Figure 12] FIG. 11 is a diagram showing an example of weight values calculated by weight calculation processing. [Figure 13] FIG. 13 is a diagram illustrating an example of setting an initial value of weights for a weighted average. [Figure 14] 13 is a flowchart illustrating an example of a weight calculation process. [Figure 15] 2 is a block diagram showing an example of a functional configuration of an information processing device; [Figure 16] 13 is a flowchart illustrating an example of a feature extraction process. [Figure 17] 2 is a block diagram showing an example of a functional configuration of an information processing device; [Figure 18] 13A and 13B are diagrams illustrating an example of display when tag information is displayed on a display unit. [Figure 19] 1 is a block diagram illustrating an example of a functional configuration of a communication system. [Figure 20] 1A to 1C are diagrams illustrating an example of transition of a display screen displayed on a display unit of an electronic device. [Figure 21] 1 is a block diagram illustrating an example of the configuration of an information processing system. [Figure 22] FIG. 1 is a diagram showing images used for learning and tags assigned as teacher labels. [Figure 23] FIG. 2 is a simplified diagram showing the contents stored in an image information DB. [Figure 24] FIG. 2 is a simplified diagram showing the contents stored in a provider information DB. [Diagram 25] FIG. 13 is a diagram showing an example of a search screen displayed on the electronic device. [Figure 26] FIG. 13 is a diagram showing an example of a selection screen displayed on the electronic device. [Figure 27]FIG. 13 is a diagram showing an example of a provider information screen displayed on the electronic device. [Figure 28] FIG. 13 is a diagram showing an example of a selection screen displayed on the electronic device. [Figure 29] 13 is a flowchart illustrating an example of a selection process. [Diagram 30] FIG. 2 is a diagram showing the contents stored in an item information DB in a simplified form. [Diagram 31] FIG. 13 is a diagram showing a display example of an editing screen displayed on the electronic device. [Diagram 32] 10 is a flowchart illustrating an example of an image editing process performed by the electronic device. [Diagram 33] FIG. 2 is a simplified diagram showing the contents stored in an image information DB. [Diagram 34] FIG. 13 is a diagram showing an example of a selection screen displayed on the electronic device. [Diagram 35] 10 is a flowchart illustrating an example of an image selection process performed by the electronic device. [Diagram 36] FIG. 13 is a diagram showing an example of a text input screen displayed on the electronic device. [Figure 37] FIG. 13 is a diagram showing an example of a selection screen displayed on the electronic device. [Figure 38] 13 is a flowchart illustrating an example of a selection process. [Figure 39] FIG. 13 is a diagram showing an example of a text input screen displayed on the electronic device. [Diagram 40] FIG. 13 is a diagram showing an example of a selection screen displayed on the electronic device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings.
[0010] [First embodiment] [Example of configuration of information processing device] 1 is a block diagram showing an example of a functional configuration of an information processing device 10. Note that the information processing device 10 can be realized by an information processing device or electronic device such as a server, a personal computer, a smartphone, or a tablet terminal.
[0011] The information processing device 10 includes an information acquisition unit 11, a tag setting unit 12, a recording control unit 13, and a storage unit 14. Each of the information acquisition unit 11, the tag setting unit 12, and the recording control unit 13 is realized by, for example, one or more processing circuits such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0012] The information acquisition unit 11 is an acquisition unit that acquires character information input by a user. For example, an input device capable of inputting character information (e.g., a keyboard, a mouse) can be used as the information acquisition unit 11. Note that a voice input device such as a microphone capable of inputting character information by voice, or an input device dedicated to voice recognition, may also be used. In addition, for example, an input device such as an imaging device may be used that is configured with a camera capable of capturing an image of character information to acquire character information, or a microphone capable of inputting voice.
[0013] When the information acquiring unit 11 receives character information input by a user, it outputs the character information to the tag setting unit 12. This character information is used when setting a reference item (tag) when extracting an image feature (feature amount). Note that in this embodiment, a feature that has been quantified will be described as a feature amount as an example of an image feature.
[0014] The tag setting unit 12 generates a list of items to be tagged (tag list) based on the character information output from the information acquisition unit 11, and outputs the character information and information on the generated tag list to the recording control unit 13. In this embodiment, an item that is a reference when extracting the features of an image is referred to as a tag. Information in which a tag and a corresponding feature amount are associated with each other is referred to as tag information. However, the tag information can also be referred to as a tag dictionary, metadata, accompanying information, additional information, etc. In this embodiment, an example in which a quantified feature amount is used as a feature corresponding to a tag is described. An example in which a score in a predetermined range (0 to 1) is used as the feature amount is shown. A method for setting multiple tags will be described in detail with reference to FIG. 5.
[0015] The recording control unit 13 executes recording control to associate the character information received by the information acquiring unit 11 with the tag list output from the tag setting unit 12 and record the associated information in the tag DB (DataBase) 100. For example, as shown in Fig. 3(A), the character information "Clothing" received by the information acquiring unit 11 and the tag list "Casual, Formal, Street, Business, Elegant, Vintage, Simple, Modern, Gothic, Feminine, Bohemian, Ethnic, Military, Rock, Punk" set by the tag setting unit 12 are stored in the tag DB 100 in association with each other.
[0016] The storage unit 14 is a storage medium that stores various information. For example, the storage unit 14 stores various information (e.g., control programs, tag DB 100) required for the information acquisition unit 11, the tag setting unit 12, and the recording control unit 13 to perform various processes. As the storage unit 14, for example, various storage media such as a ROM (Read Only Memory), a RAM (Random Access Memory), an SRAM (Static Random Access Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a combination of these can be used.
[0017] The tag DB 100 is a database that stores character information input by a user and a list of tags related to the character information. The tag list and the like stored in the tag DB 100 will be described in detail with reference to FIG.
[0018] 1 shows an example in which the information acquisition unit 11 and the storage unit 14 are provided in the information processing device 10, but at least one of them may be used as a separate device different from the information processing device 10. For example, by registering the information acquisition device and the storage device in advance in the information processing device 10 (for example, by pairing or connecting using a wired line or a wireless line), it is possible to make them function as the information acquisition unit and the storage unit of the information processing device 10.
[0019] [Example of configuration of information processing device] 2 is a block diagram showing an example of a functional configuration of the information processing device 50. Note that the information processing device 50 can be realized by an information processing device or electronic device such as a server, a personal computer, a smartphone, or a tablet terminal.
[0020] The information processing device 50 includes an image acquisition unit 51, a text information generation unit 52, a feature extraction unit 53, a recording control unit 54, and a storage unit 55. Each of the image acquisition unit 51, the text information generation unit 52, the feature extraction unit 53, and the recording control unit 54 is realized by, for example, one or more processing circuits such as a CPU, a GPU, etc. Also, although the information processing device 10 and the information processing device 50 are shown as separate entities in Figs. 1 and 2, they may be configured as an integrated device. An example of such an integrated device is shown in Fig. 17.
[0021] The image acquisition unit 51 is an acquisition unit that acquires image information input by a user. For example, an input device (for example, a recording medium reader, a camera) capable of inputting image information can be used as the image acquisition unit 51. In this embodiment, when an image or image information is referred to, it means both the image and the image file corresponding to the image, or either one of them. For example, the image 41 can be read and acquired from a recording medium (for example, a memory card, a Universal Serial Bus (USB) memory, a Hard Disk Drive (HDD), a Compact Disc (CD), a Digital Versatile Disc (DVD), a Blu-ray (registered trademark) Disc (BD), etc.) that stores an image file in which a plurality of image groups 40 (including the image 41) are stored. Also, for example, the recording medium in which the image 41 is stored and the information processing device 50 can be connected using wireless communication or wired communication, and the image acquisition unit 51 can read and acquire the image 41 from the recording medium. Also, when the recording medium is built into the information processing device 50, the image acquisition unit 51 can read and acquire the image 41 from the recording medium. Alternatively, image 41 may be acquired using an imaging device. For example, an image of a subject included in image 41 may be captured by the imaging device. In this manner, image acquisition unit 51 is realized by an input interface, a file input device, an imaging device, etc.
[0022] The text information generating unit 52 generates text information based on the image information output from the image acquiring unit 51. Then, the text information generating unit 52 outputs the generated text information to the feature extracting unit 53. This text information generating process will be described in detail with reference to FIG.
[0023] The feature extraction unit 53 extracts a plurality of feature amounts (feature amounts related to text information) corresponding to a plurality of tags stored in the tag DB 100 based on the text information output from the text information generation unit 52. Then, the feature extraction unit 53 outputs the extracted feature amounts of the plurality of tags to the recording control unit 54. The extraction process of the feature amounts of the plurality of tags will be described in detail with reference to FIG.
[0024] The recording control unit 54 executes recording control for associating the multiple feature amounts for each of the multiple tags extracted by the feature extraction unit 53 with the image information acquired by the image acquisition unit 51 and recording them in the image DB 200. For example, they are stored in the image DB 200 as shown in FIG.
[0025] The storage unit 55 is a storage medium that stores various information. For example, the storage unit 55 stores various information (e.g., a control program, a tag DB 100, and an image DB 200) required for the image acquisition unit 51, the text information generation unit 52, the feature extraction unit 53, and the recording control unit 54 to perform various processes. As the storage unit 55, for example, various storage media such as a ROM, a RAM, an SRAM, an HDD, an SSD, or a combination of these can be used. The tag DB 100 corresponds to the tag DB 100 shown in FIG. 1.
[0026] 2 shows an example in which the image acquisition unit 51 and the storage unit 55 are provided in the information processing device 50, but at least one of them may be used as a separate device different from the information processing device 50. For example, by registering the image acquisition device and the storage device in the information processing device 50 in advance, it is possible to make them function as the image acquisition unit and the storage unit of the information processing device 50.
[0027] The image DB 200 is a database that stores the feature amounts extracted by the feature extraction unit 53 and the image information acquired by the image acquisition unit 51 in association with each other. The image DB 200 will be described in detail with reference to FIG.
[0028] [Tag DB configuration example] FIG. 3A is a simplified diagram showing the contents of a list of objects and tags stored in the tag DB 100. As shown in FIG.
[0029] The tag DB 100 is a database for storing a plurality of tags set by the tag setting unit 12 (see FIG. 1).
[0030] Specifically, an object 101 acquired by an information acquisition unit 11 (see FIG. 1) and a tag list 102 indicating a list of a plurality of tags set by a tag setting unit 12 are stored in a tag DB 100 in association with each other.
[0031] As will be described later, it is possible to generate a list of tags using preset information (for example, a database). An example of the information used in this case is shown in FIG.
[0032] FIG. 3B is a simplified diagram showing the contents of the tag-related information stored in the tag DB 100a.
[0033] The tag DB 100a is a database for storing information (target 101, tag list 102) used when the tag setting unit 12 sets multiple tags, and information (selection information 103) about the multiple tags set by the tag setting unit 12. A method for setting multiple tags will be described in detail with reference to FIG.
[0034] [Image DB configuration example] FIG. 4 is a simplified diagram showing the contents of image information and tag information stored in the image DB 200. As shown in FIG.
[0035] The image DB 200 is a database for storing images acquired by the image acquisition unit 51 and feature amounts of a plurality of tags extracted for each tag by the feature extraction unit 53 (see FIG. 2) in association with each other.
[0036] Specifically, an image 201 acquired by the image acquisition unit 51, a tag 202 whose features have been extracted by the feature extraction unit 53, and a plurality of features 203 extracted for each tag are stored in the image DB 200 in association with each other.
[0037] [Example of generating a tag list] Next, a setting method in which the tag setting unit 12 sets a plurality of tags desired by the user U1 using character information input by the user U1 will be described.
[0038] FIG. 5 is a diagram showing the flow of a setting method for setting a plurality of tags desired by user U1 using character information input by user U1.
[0039] FIG. 5(A) shows an example in which user U1 inputs the characters "Clothes". For example, a display screen 16 for inputting character information can be displayed on the input / output device 15, and the characters "Clothes" can be input using the display screen 16. The input / output device 15 may be provided in the information processing device 10, or an external device different from the information processing device 10 may be used. The input / output device 15 may utilize a user interface such as a touch panel, or may be used together with other operating members (for example, a keyboard, a mouse). The object shown here is used as a criterion when setting a tag, and includes the meaning of, for example, a category, a classification, etc.
[0040] When user U1 inputs the characters “Clothes” into character input field 17 on display screen 16 and selects decision button 18, information acquisition unit 11 accepts character information regarding the characters “Clothes” input into character input field 17 and outputs the character information to tag setting unit 12.
[0041] FIG. 5(B) shows a list of tags set by the tag setting unit 12. For example, the tag setting unit 12 can generate a list of tags using an AI (Artificial Intelligent) model (for example, a machine learning model generated by machine learning). Note that the learning shown in this embodiment means finding the regularity behind a large amount of data based on the data. Also, the AI model generated by learning shown in this embodiment is generated by various learning algorithms. For example, various images and sentences describing each of these images are read as teacher data and learned in advance, and this AI model can be used for text information generation processing, etc.
[0042] As the AI model, for example, a large language model (LLM) can be used. As the LLM, for example, ChatGPT (Generative Pre-trained Transformer), Bard, Llama (Large Language Model Meta AI), Gemini, Claude, etc. can be used. Note that these are only examples, and other AI models may be used. For example, an object (for example, clothing) specified by a user and condition information for outputting a plurality of tags capable of providing features related to a specific viewpoint (for example, the viewpoint of a clothing expert) corresponding to the object can be input to the LLM as instruction information (prompt), and the output information in response to this can be a tag list. Here, the specific viewpoint is, for example, a viewpoint that can be determined according to the sensibility of a person of a certain age with a certain attribute (for example, gender, occupation, career), and can also be referred to as the viewpoint when the person views the object.
[0043] Here, an example of a prompt that suggests tags using ChatGPT is shown. For example, the characters "Clothing" are entered in the character input field 17 as "category information (e.g., product category)". In addition, as an expert on the entered "category information" (e.g., clothing), an instruction (command) is given to execute an operation to display a list of tags that satisfy the "tag list requirements" for the entered "category information" in an "output format". Here, it is possible to give an instruction (command) to satisfy, for example, that the "tag list requirements" are words that represent the style or design of the "category information (e.g., product category)".
[0044] Also, the tag setting unit 12 can generate a list of tags using, for example, information that has been set in advance. For example, as shown in FIG. 3B, the tag list can be generated using a tag DB 100a that associates an object 101 with a tag list 102. For example, the information acquisition unit 11 is a reception unit that receives a selection operation in which a user selects a character string corresponding to the object 101 in the tag DB 100a. In this case, the tag setting unit 12 can set the tag list 102 associated with the object 101 corresponding to the character string received by the information acquisition unit 11. For example, when "clothing" is received by the information acquisition unit 11, the tag list 102 "casual, formal, ..., rock, punk" associated with the object 101 "clothing" can be set. In this case, selection information indicating that the selection has been made is given to the selection information 103 of the tag DB 100a. Note that FIG. 3B shows an example in which "1" is stored as the selection information indicating that the selection has been made.
[0045] [Example of generating a tag list using spatial and generational language differences] Here, it is assumed that a plurality of languages are used for the characters acquired by the information acquisition unit 11. For example, it is assumed that Japanese, English, and other characters are input for a Japanese person. Therefore, it is possible to estimate the regionality of the user who performed the input operation based on the input language. It is also assumed that different words are used for the characters acquired by the information acquisition unit 11 depending on the generation. For example, words used only by the younger generation, words only understood by the elderly generation, and the like are assumed. When these words are input, it is possible to estimate the generation of the user. For example, when the Japanese word "AB (where AB is Japanese)" which is assumed to be frequently used by women in their twenties is input, it is possible to estimate the attributes of the user as "Japanese", "twenties", and "female". Also, for example, when the Japanese word "CD (where CD is Japanese)" which is assumed to be frequently used by men in their seventies is input, it is possible to estimate the attributes of the user as "Japanese", "elderly", and "male". In this way, since it is possible to estimate the attributes of the user based on the characters acquired by the information acquisition unit 11, tags may be set in consideration of the attributes of the user. For example, it is possible to use LLM to input words used by a user and output the user's attributes. Therefore, it is possible to set tags that take into account the user's attributes that can be estimated based on the words used by the user using LLM. This makes it possible to set appropriate tags that take into account spatial and generational language differences.
[0046] Furthermore, the tag setting unit 12 outputs tag list information relating to the multiple tags that have been set to the recording control unit 13. Then, the recording control unit 13 stores multiple tags corresponding to the tag list information output from the tag setting unit 12 in the tag DB 100, as shown in Fig. 5(B). Note that Fig. 5 shows an example in which "casual, formal, ..., rock, punk" are determined as tags corresponding to the characters "Western clothing" input by the user. Furthermore, these contents correspond to Fig. 3(A).
[0047] In FIG. 5, an example is shown in which multiple tags are set based on the character "Clothes" input by the user U1, but other setting methods may be used. For example, the tag setting unit 12 determines other categories related to "Clothes" (for example, categories considering the era, regionality, and appearance) based on the character information "Clothes" output from the information acquisition unit 11. For example, a category DB that stores a large number of character information (for example, clothes, hats, shoes, bags, socks), categories to which items corresponding to these character information belong (for example, clothing, accessories, fashion), and multiple tags corresponding to the categories (for example, "casual, formal, ..., rock, punk") is generated in advance, and a category can be determined based on this category DB. In addition, a category related to the input character information may be determined using other AI technology. Then, the tag setting unit 12 can extract multiple tags corresponding to the determined category from the category DB and make the multiple tags into a tag list.
[0048] [Example of feature extraction for each tag] Next, a method will be described in which the information processing device 50 uses image information input by the user U1 to extract a feature amount for each tag related to the image information.
[0049] FIG. 6 is a diagram showing the flow of an extraction process for extracting a feature amount for each tag related to the image 41. In FIG.
[0050] 6A shows an image 41 (an image to be tagged) input by a user U1. The image 41 is acquired by an image acquisition unit 51.
[0051] FIG. 6B shows text information 42 generated by the text information generating unit 52 based on the image 41. When the image 41 acquired by the image acquiring unit 51 is input to the text information generating unit 52, the text information generating unit 52 generates the text information 42 based on the image 41. The text information shown here is, for example, text formed by characters such as words and sentences written in a predetermined language that can be handled by humans. Specifically, the text information generating unit 52 generates text for each object included in the image 41 (for example, a subject and its background in the case of a captured image, a model image (or motif image) and its background image in the case of a generated image, etc.), a description of the object itself (for example, appearance, form, color, movement, impression, attribute (for example, properties or characteristics belonging to the object itself)), a description of the relationship with other objects, a description of an event caused by the interaction of each object, etc., and generates the text information 42. That is, the text information generating unit 52 generates text describing one or more objects included in the image 41 as the text information 42. It should be noted that the text information can be generated using a known text generation method.
[0052] For example, the text information generating unit 52 can use a database prepared in advance to generate text information 42 based on the image 41. For example, the text information generating unit 52 can detect one or more objects (and their attributes, sizes, positions in the image, etc.) included in the image 41, and generate a sentence based on each piece of information by linking descriptions of those objects according to a predetermined rule.
[0053] Also, for example, the text information generating unit 52 can generate the text information 42 based on the image 41 by using an AI model. As this AI model, for example, image to text and img2txt can be used. Note that these can be executed by installing a predetermined application in the information processing device 50.
[0054] Also, it is possible to generate text information using other AI models. For example, it is possible to use a multimodal LLM. The multimodal LLM is composed of an image processing part (pre-trained part) used for image recognition processing, a large-scale language model part, and an adapter part connecting them. Then, when an image 41 acquired by the image acquisition unit 51 is input to the multimodal LLM, it is possible to output text information related to the image 41. Note that when an LLM such as the multimodal LLM is used, the amount of calculation is large, so it is expected that the processing load will be large. Therefore, for example, when the information processing device 50 is a device with a slow calculation speed such as a personal computer, a device with a fast calculation speed (for example, a server) may be used. For example, the text information generation unit 52 transmits the image 41 acquired by the image acquisition unit 51 to a server, executes a process of generating text information in this server, and the text information generation unit 52 can receive and use the processing result (text information) from the server.
[0055] Also, text information may be generated using an object designated by a user, multiple tags set by the tag setting unit 12, and the like. For example, when using a multimodal LLM, it is possible to generate text information taking into consideration an object designated by a user (for example, clothing). For example, condition information for outputting a sentence capable of providing a specific viewpoint corresponding to an object designated by a user (for example, clothing) as text information related to an image 41 acquired by the image acquisition unit 51 can be input to the LLM as instruction information (prompt), and output information in response to this can be made text information. Also, for example, text information can be generated taking into consideration multiple tags (for example, casual, ..., punk) set by the tag setting unit 12. For example, instruction information (prompt) for generating a sentence related to multiple tags (for example, casual, ..., punk) can be input to the multimodal LLM, and output information in response to this can be made text information. Also, by inputting both of these instruction information (prompts) to the multimodal LLM and making output information in response to this text information, it is possible to generate text information taking into consideration an object designated by a user (for example, clothing) and multiple tags (for example, casual, ..., punk).
[0056] Here, for example, when a human sees an object (e.g., image 41), it is easy to put into words the superficial emotions of the object (e.g., cute, fun, sad). In other words, superficial emotions are easy for humans to express in words. On the other hand, complex emotions based on various memories born from humans' past experiences (e.g., intrinsic emotions, subtle emotions), etc., are often difficult to put into words. Such intrinsic emotions and oozing emotions that humans cannot put into words can be expressed by converting the emotions into text based on the image, and the expressed emotions can be used as analysis targets for extracting the features of multiple tags. In other words, when extracting the features of multiple tags related to an image, by using text information that is converted from the viewpoint, emotion, image, etc. of a certain person, it is possible to extract each feature based on limited content for the image. This makes it possible to extract appropriate features from a certain viewpoint for multiple tags corresponding to that viewpoint.
[0057] FIG. 6C shows tag information 43 generated based on the feature extraction unit 53. When the text information 42 generated by the text information generation unit 52 is input to the feature extraction unit 53, the feature extraction unit 53 extracts feature amounts for each of the multiple tags stored in the tag DB 100 based on the text information 42. The feature amount shown here means a value (score) in a predetermined range indicating the relevance (or relationship, correlation) between what is included in the text information 42 (for example, one or multiple objects and their backgrounds) and the tag. In this embodiment, an example is shown in which 0 to 1 is used as the value in the predetermined range. Also, it is assumed that the value becomes larger (maximum value 1) as the relevance with the tag becomes higher, and the value becomes smaller (minimum value 0) as the relevance with the tag becomes lower. For example, the value "0.4" of the tag "casual" indicates a value related to "casual", and the value "0.7" of the tag "formal" indicates a value related to "formal". For example, since the value "0.7" of the tag "formal" is a relatively high value, it is estimated that the relevance between the object (one piece) included in the text information 42 and "formal" is high. On the other hand, for example, the numeric value "0.0" of the tag "rock" indicates a numeric value related to "rock," and the numeric value "0.0" of the tag "punk" indicates a numeric value related to "punk." Since these values are 0, it is presumed that there is almost no correlation between the object (one piece) included in the text information 42 and "rock" and "punk." Note that a publicly known method for extracting feature quantities can be used as a method for extracting feature quantities.
[0058] For example, the feature extraction unit 53 can extract features for each of a plurality of tags based on the text information 42 using a database prepared in advance. For example, the feature extraction unit 53 detects, for each tag, one or more character strings that are related to the character example of each tag among the character strings included in the text information 42, using the database. Then, the feature extraction unit 53 can calculate a score for each tag based on the number of character strings detected for each tag, the contents of the character strings, etc. For example, in the case of a tag with a large number of detected character strings, it is possible to calculate a high score according to the number. Also, for example, in the case of a tag with important contents of the detected character strings, it is possible to calculate a high score according to the importance of the detected character strings.
[0059] In addition, for example, artificial intelligence (AI) or the like can be used to extract feature amounts of multiple tags for each tag based on the text information 42. For example, an AI model can be used to extract feature amounts of multiple tags for each tag. For example, the above-mentioned LLM can be used as this AI model.
[0060] For example, the feature extraction unit 53 can use the LLM to extract the feature amount of a plurality of tags for each tag based on the text information 42. As described above, for example, ChatGPT, Bard, Llama, Gemini, Claude, etc. can be used as the LLM. Note that these are only examples, and other models may be used. For example, it is possible to input instruction information (prompt) to the LLM to output a score (0 to 1) regarding the relevance with each of the plurality of tags using a sentence constituting the text information 42, and to use the output information as the feature amount for each of the plurality of tags. Also, for example, it is possible to input condition information (prompt) to the LLM to output a score (0 to 1) regarding the relevance with each of the plurality of tags from a specific perspective (for example, the perspective of a clothing expert) corresponding to an object (for example, clothing) specified by the user using a sentence constituting the text information 42, and to use the output information as the feature amount for each of the plurality of tags.
[0061] In addition, when LLM is used, it is expected that the processing load will be large due to the large amount of calculation. Therefore, for example, when the information processing device 50 is a device with a slow calculation speed such as a personal computer, a device with a fast calculation speed (for example, a server) may be used. For example, the feature extraction unit 53 can transmit the text information 42 output from the text information generation unit 52 to a server, execute a process of extracting the feature amount for each of the multiple tags in this server, and receive the processing result (the feature amount for each of the multiple tags) from the server and use it.
[0062] Here, an example of a prompt for tagging a sentence using ChatGPT is shown. For example, the characters "Clothing" entered in the character input field 17 are entered as "category information", and the text information 42 generated by the text information generating unit 52 is entered as "text information". In addition, as an expert on the entered "category information" (e.g. clothing), an instruction (command) is given to receive the entered "text information", convert the contents of the entered "text information" into a "vector" representing the style and atmosphere of the entered "category information", and execute an operation to display the "vector" in an "output format". Here, as a definition of the "vector", each element of the "vector" is a real number between 0 and 1, and an instruction (command) can be given to indicate that the "vector" has elements of multiple tags "casual, ..., punk" set by the tag setting unit 12.
[0063] FIG. 6(D) shows the relationship between the image 41 acquired by the image acquisition unit 51 and the tag information 43 generated by the feature extraction unit 53. When the tag information 43 generated by the feature extraction unit 53 is input to the recording control unit 54, the recording control unit 54 associates the image 41 output from the image acquisition unit 51 with the tag information 43 and stores them in the image DB 200. For example, as shown in FIG. 4, the image 41 and the tag information 43 are stored in association with each other. Note that FIG. 6 shows an example in which the image 41 and the tag information 43 are associated with each other and stored in the image DB 200, but the image 41 and the tag information 43 may be associated with each other and output from an output device or output to another device. For example, as shown in FIG. 18 and FIG. 20, the image 41 and the tag information 43 can be associated with each other and displayed on the display units 821 and 931. This allows the user U1 to easily grasp the feature amounts of multiple tags related to the target specified by the user U1 for the image 41. In the above, an example has been shown in which multiple tags are set and features are extracted for each of these tags, but this embodiment can also be applied to a case in which one tag is set and features are extracted for that tag.
[0064] [Example of operation of information processing device] Fig. 7 is a flowchart showing an example of a tag setting process in the information processing device 10. This tag setting process is executed based on a program stored in the storage unit 14. This tag setting process is executed when a user operation is accepted by the information acquisition unit 11. This tag setting process will be described with reference to Figs. 1 to 6 as appropriate.
[0065] In step S401, the information acquiring unit 11 accepts characters entered by the user U1, and outputs character information relating to the accepted characters to the tag setting unit 12.
[0066] In step S402, the tag setting unit 12 sets a tag list based on the character information output from the information acquisition unit 11, and outputs the set tag list to the recording control unit 13. The setting method is similar to the example shown in FIG.
[0067] In step S403, the recording control unit 13 stores the tag list set by the tag setting unit 12 in the tag DB 100. For example, the tag list is stored in the tag DB 100 as shown in FIG.
[0068] [Example of operation of information processing device] Fig. 8 is a flowchart showing an example of feature extraction processing in the information processing device 50. This feature extraction processing is executed based on a program stored in the storage unit 55. This feature extraction processing is executed when an image is acquired by the image acquisition unit 51. This feature extraction processing will be described with appropriate reference to Figs. 1 to 7.
[0069] In step S411, the image acquisition unit 51 acquires an image input by the user U1, and outputs image information relating to the acquired image to the text information generation unit 52.
[0070] In step S412, text information generation unit 52 generates text information based on the image information output from image acquisition unit 51, and outputs the generated text information to feature extraction unit 53. The method of generating this text information is similar to the example shown in FIG.
[0071] In step S413, the feature extraction unit 53 extracts feature amounts for each of the multiple tags stored in the tag DB 100 based on the text information generated by the text information generation unit 52. Then, the feature extraction unit 53 outputs the extracted feature amounts for each of the multiple tags to the recording control unit 54. The method of extracting the feature amounts for each tag is the same as the example shown in FIG.
[0072] In step S414, the recording control unit 54 stores the feature amounts for each of the multiple tags output from the feature extraction unit 53 and the image information output from the image acquisition unit 51 in the image DB 200. For example, as shown in Fig. 4, the feature amounts for each of the multiple tags and the image information are stored in the image DB 200 in association with each other.
[0073] [Example of generating text information in multiple languages] Here, when expressing the same image, the expression differs depending on the language, and the impression that humans get is often different. For example, when expressing the same image in Japanese and English, the expression differs between Japanese and English, and the impression that humans get is also considered to be different. In addition, for example, a case is assumed in which the text information generating unit 52 can generate text information in a plurality of languages. For example, it is possible to generate text information in each language such as Japanese, English, French, and Spanish. In this case, the text information generating unit 52 generates text information in a plurality of languages based on the image 41, and the feature extracting unit 53 may extract the feature amount for each of the plurality of tags stored in the tag DB 100 based on the text information in the plurality of languages generated by the text information generating unit 52. In this case, for example, it is possible to input instruction information (prompt) to output scores (feature amount) for the plurality of tags to the LLM using each sentence constituting the text information in the plurality of languages, and to set the output information in response to this as the feature amount for each of the plurality of tags.
[0074] [Example of tagging based on image] In the above, an example has been shown in which multiple tags are set based on an object designated by the user U1, and feature quantities related to the multiple tags are extracted. However, multiple tags may be set based on some information related to an image, and feature quantities related to the multiple tags may be extracted. For example, an object detection process may be performed on the image 41 acquired by the image acquisition unit 51, and a tag may be set based on the detected object. For example, since the image 41 includes a dress, the dress is detected from the image 41 by the object detection process. In addition, the attributes of the detected dress can be detected by the image recognition process. For example, the color, shape, size, etc. related to the dress can be detected. Therefore, it is possible to determine what kind of clothing the detected dress is. For example, if the image 41 includes a green long-sleeved dress with an A-line silhouette and a round neck, it is possible to estimate that a tag related to "clothing" is set for the image 41.
[0075] For example, if an image contains a sofa, a chair, and a table, these can be detected from the image by object detection processing. Also, attributes of the detected objects (e.g., size, color, position) can be detected by image recognition processing. In this way, if an image contains a sofa, a chair, and a table, it can be estimated that a tag related to "furniture" is set for the image.
[0076] Also, for example, if an image includes a house, the house can be detected from the image by object detection processing. Also, attributes of the detected house (e.g., size, color, position) can be detected by image recognition processing. In this way, if a house is included in an image, it can be estimated that a tag related to "house" is set for the image. Tag setting based on these object detections can be realized by including an object detection unit in the tag setting unit 12 (see FIG. 17).
[0077] [Example of setting tags based on text information] Here, an example is shown in which multiple tags are set based on text information generated from an image, and feature quantities related to the multiple tags are extracted. For example, character recognition processing may be performed on the text information 42 generated by the text information generating unit 52, and tags may be set based on character strings, words, and the like included in the text information 42. For example, since the image 41 includes a dress, it is estimated that the text information 42 also includes a description related to the dress. It is also estimated that the text information 42 includes a description related to the attributes of the dress. In this case, the description related to the dress and the description related to the attributes of the dress included in the text information 42 can be detected by character recognition processing. Based on each of these pieces of information, multiple tags can be set in the same way as when tags are set based on an image. For example, if the text information 42 includes a description related to a green long-sleeved dress with an A-line silhouette and a round neck, it is possible to set a tag related to "clothing" for the image 41.
[0078] [Example of feature extraction using one multimodal LLM] Here, it is also possible for one multimodal LLM to handle both the text information generation process by the text information generation unit 52 and the feature extraction process by the feature extraction unit 53. Thus, an image may be input to the multimodal LLM, and a prompt for generating features for each of a plurality of tags from the image may be input to extract features related to the image. In this case, both the text information generation process and the feature extraction process are executed in the multimodal LLM, and features related to the image are extracted based on the results of each process. That is, it is possible for the multimodal LLM to function as the text information generation unit 52 and the feature extraction unit 53.
[0079] As a prompt when using the multimodal LLM, it is possible to input a prompt that simultaneously generates text and extracts features for each tag. For example, the characters "Clothing" input in the character input field 17 are input as "category information", and the image 41 acquired by the image acquisition unit 51 is input as "input image". In addition, as an expert on the input "category information" (e.g. clothing), an instruction (command) is given to receive the input "input image", convert the object reflected in the input "input image" into a "vector" representing the style and atmosphere of the input "category information", and display the "vector" in an "output format". Here, as a definition of the "vector", it is possible to give an instruction (command) that each element of the "vector" is a real number between 0 and 1, and that the "vector" has elements of multiple tags "casual, ..., punk" set by the tag setting unit 12.
[0080] [Example of effect of the first embodiment] In this way, according to this embodiment, when extracting feature amounts related to multiple tags for a subject included in an image, it is possible to extract feature amounts of multiple tags for each tag based on text information generated from the image. That is, it is possible to extract feature amounts of multiple tags related to an image for each tag using the text information generating unit 52 that generates text information based on an image and the feature extracting unit 53 that extracts feature amounts of multiple tags for each tag based on the text information. For example, it is possible to use a general-purpose AI model for both the text information generating unit 52 and the feature extracting unit 53. In this case, a large amount of teacher data for extracting the feature amount of an image is not required. That is, it is possible to appropriately extract the feature amount of an image without preparing a large amount of teacher data. In addition, since the feature amount of multiple tags is extracted for each tag based on text information generated from an image, it is possible to extract new feature amounts different from the feature amount directly extracted from an image. That is, it is possible to extract new feature amounts in which differences caused by differences between languages are taken into consideration. In this way, according to the first embodiment, it is possible to appropriately extract the feature of an image.
[0081] In addition, for example, by using LLM, multimodal LLM, etc., it is possible to improve the processing speed of the extraction process of feature amounts related to multiple images. In other words, it is possible to make the extraction process of feature amounts related to multiple images more efficient. In addition, by using LLM, multimodal LLM, etc., it is possible to significantly reduce the amount of manual work. In other words, it is possible to achieve significant labor savings.
[0082] [Second embodiment] In the first embodiment, an example is shown in which tag information is generated using one type of text information generating unit 52 and one type of feature extracting unit 53. However, tag information may be generated using two or more types of text information generating units or two or more types of feature extracting units. Alternatively, tag information may be generated using both two or more types of text information generating units and two or more types of feature extracting units. Thus, FIG. 9 and other figures show an example in which tag information is generated using two or more types of text information generating units and one type of feature extracting unit.
[0083] [Example of configuration of information processing device] Fig. 9 is a block diagram showing an example of a functional configuration of the information processing device 500. The information processing device 500 is a modified version of the information processing device 50 shown in Fig. 2, and is common to the information processing device 50 except that the image DB 200 is omitted, an acquisition unit 510, a text information generation unit 520 and a feature extraction unit 530 are provided instead of the image acquisition unit 51, the text information generation unit 52 and the feature extraction unit 53, and a weight calculation unit 540 and a weight DB 600 are added. For this reason, the parts common to the information processing device 50 are given the same reference numerals as those of the information processing device 50, and some of the descriptions thereof will be omitted.
[0084] The information processing device 500 includes an acquisition unit 510, a text information generation unit 520, a feature extraction unit 530, a weight calculation unit 540, a recording control unit 54, and a storage unit 55. Each of the acquisition unit 510, the text information generation unit 520, the feature extraction unit 530, the weight calculation unit 540, and the recording control unit 54 is realized by, for example, one or more processing circuits such as a CPU, a GPU, etc. Also, like the example shown in Fig. 2, the information processing device 10 and the information processing device 500 are shown as separate entities, but the two may be configured as an integrated device.
[0085] The acquisition unit 510 is an acquisition unit that acquires tagged data TD1 associated with image information TD2 and tag information TD3. Then, the acquisition unit 510 outputs the image information TD2 to the text information generation unit 520, and outputs the tag information TD3 to the weight calculation unit 540. For example, the acquisition unit 510 can acquire an image file associated with the image information TD2 and the tag information TD3. The tagged data TD1 is data used when determining a weight (weight value), and corresponds to, for example, teacher data of a neural network. However, in the second embodiment, an example is shown in which only one or a small amount (a predetermined number or less) of teacher data is used. For example, the number of tagged data TD1 can be several times to several tens of times the number of generation units constituting the text information generation unit 520 (i.e., the number of weights (the number of models)).
[0086] The text information generating unit 520 includes a plurality of generating units (first generating unit 521 to Nth generating unit 524). Each of these generating units (first generating unit 521 to Nth generating unit 524) generates text information based on the image information TD2 output from the acquiring unit 510. Then, the text information generating unit 520 outputs the text information generated by each generating unit to the feature extracting unit 530. Here, N is an integer of 2 or more. That is, in FIG. 9, four generating units (first generating unit 521 to Nth generating unit 524) are illustrated as an example, but two, three, or five or more generating units may be included. Also, the first generating unit 521 to Nth generating unit 524 execute a process for converting an image into text information by different algorithms. In this way, when generating text information based on the same image, a plurality of pieces of text information are generated by using different algorithms, and thus a plurality of pieces of text information each having different sentences are generated. By using the text information of these different sentences to extract the feature amount for each tag, it becomes possible to extract contents that cannot be extracted by one piece of text information.
[0087] As described above, it is also possible to generate text information in multiple languages based on one image. Therefore, the first generating unit 521 to the Nth generating unit 524 may generate text information in multiple languages based on one image. In this way, by generating text information in multiple languages based on the same image, multiple pieces of text information, each of which is a different sentence, are generated.
[0088] As described above, in the first embodiment, one model of the text information generation unit 52 (e.g., img2txt) is used, and therefore it is considered that there is a large dependency on the characteristics of that model. In contrast, in the second embodiment, multiple models of text information generation units (e.g., img2txt) are used. This makes it possible to reduce the dependency on the characteristics of one model. Furthermore, by using multiple models in this way, it becomes possible to stabilize behavior. The process of generating text information using multiple generation units will be described in detail with reference to FIG. 11 etc.
[0089] The feature extraction unit 530 extracts feature amounts for each of the plurality of pieces of text information for each tag, based on the tag list stored in the tag DB 100 and the plurality of pieces of text information output from the text information generation unit 520. Then, the feature extraction unit 530 outputs each of the extracted feature amounts to the weight calculation unit 540. This feature amount extraction process will be described in detail with reference to FIG. 11 etc.
[0090] The weight calculation unit 540 calculates weights to be used by the weighted average calculation unit 550 (see FIG. 15) for each of the multiple generation units (first generation unit 521 to Nth generation unit 524) based on the feature amount for each tag related to each of the multiple text information output from the feature extraction unit 530 and the tag information TD3 output from the acquisition unit 510. Then, the weight calculation unit 540 outputs the calculated weight values to the recording control unit 54. The calculation process of the weights will be described in detail with reference to FIG. 13 and the like.
[0091] The memory unit 55 stores, for example, various information (e.g., control programs, tag DB 100, weight DB 600) required by the acquisition unit 510, text information generation unit 520, feature extraction unit 530, weight calculation unit 540, and recording control unit 54 to perform various processing.
[0092] The weight DB 600 is a database that stores the weight values calculated by the weight calculation unit 540 in association with a plurality of generating units (first generating unit 521 to N-th generating unit 524). The weight DB 600 will be described in detail with reference to FIG.
[0093] [Example of weight DB configuration] Fig. 10 is a simplified diagram showing the contents of weight data stored in weight DB 600. Weight DB 600 is a database for storing weight values used when weighted average calculation section 550 (see Fig. 15) executes weighted average calculation processing.
[0094] Specifically, generation unit identification information 601 and weight 602 are stored in association with each other in a weight DB 600 .
[0095] The generation unit identification information 601 stores information for identifying a plurality of generation units (first generation unit 521 to Nth generation unit 524) constituting the text information generation unit 520. For ease of explanation, only the names corresponding to the respective generation units are shown in FIG.
[0096] The weight 602 stores the weight values calculated by the weight calculation unit 540. Here, the total value of each weight (α1+α2+α3+...+α N ) is set to 1. In addition, the weight α i A specific method for calculating is shown in Figures 11 and 13. Note that i is an integer that satisfies 0≦i≦N.
[0097] [Example of weight calculation process] Fig. 11 is a diagram showing a schematic flow of the weight calculation process by the information processing device 500. Fig. 12 is a diagram showing an example of weight values calculated by the weight calculation process by the information processing device 500.
[0098] 11, when image information TD2 constituting tagged data TD1 is input to text information generator 520, first generator 521 to Nth generator 524 generate text information TX1 to TX4 based on image information TD2. Note that, in order to facilitate explanation, in FIG. 11, of the text information TX1 to TX4, only the sentences related to the text information TX1 and TX2 are shown in a simplified manner, and the sentences related to the text information TX3 and TX4 are omitted.
[0099] Next, when the text information TX1 to TX4 generated by the first generation unit 521 to the Nth generation unit 524 are input to the feature extraction unit 530, the feature extraction unit 530 generates tag information FQ1 to FQ4 based on the text information TX1 to TX4. In addition, in Fig. 11, for ease of explanation, the tag information FQ1 to FQ4 is shown in a simplified manner within a rectangle indicating the feature extraction unit 530. Also, of the tag information FQ1 to FQ4, only data related to the tag information FQ1 and FQ2 is shown in a simplified manner, and data related to the tag information FQ3 and FQ4 is omitted.
[0100] Next, the tag information FQ1 to FQ4 generated by the feature extraction section 530 and the tag information TD3 constituting the tagged data TD1 are input to the weight calculation section 540.
[0101] Next, the weight calculation unit 540 calculates the weight value (weight value for each pipeline) for each generation unit (first generation unit 521 to Nth generation unit 524) based on the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the tag information TD3 constituting the tagged data TD1. Here, a method for determining the weight of the weighted average from the loss function will be described. For example, the weight of the weighted average can be determined using the gradient descent method, similar to the method used in deep learning.
[0102] Here, a value obtained by extracting only the feature amount of tag information will be referred to as a tag value vector for explanation. Specifically, the tag value vector is a vector whose number of dimensions is the number of tags set by the tag setting unit 12. For example, in Figs. 9 to 11, an example in which the number of tags is 15 is shown, so the tag value vector is a 15-dimensional vector. Therefore, an example in which the tag value vector is a 15-dimensional vector will be shown below.
[0103] In addition, the tag value vector (output value of the feature extraction unit 530) generated for each generation unit (first generation unit 521 to Nth generation unit 524) is expressed as a vector x i (x i1 ,x i2 ,…,x i15 ) and the vector x i The weight of α i Here, i is an integer satisfying 0≦i≦N. Note that vector x i can also be referred to as a tag value vector of the output of each pipeline (first generation unit 521 to Nth generation unit 524 → feature extraction unit 530). In addition, the tag value vector output from the feature extraction unit 530 is referred to as vector y (y1, y2, ..., y 15 ) That is, the following formula 1 holds. Note that the actual vector x i , vector y is a column vector, but for ease of explanation, i , vector y is expressed as a row vector.
[0104]
number
[0105] In addition, the tag value vector of the tagged data (corresponding to the training data) is expressed as vector t(t1, t2, …, t 15 )
[0106] Next, the procedure of the method for determining the weights of the weighted average will be described.
[0107] First, the weight α iFor example, it can be determined based on the following formula 2. Note that N indicates the number of pipelines. α i =1 / N Equation 2
[0108] Next, the loss function L is defined. Here, the loss function L is defined based on the following Equation 3. L = |vector t - vector y| 2 formula 3
[0109] Next, for each α of the loss function L i The differentiation with respect to is calculated using the following equation 4.
[0110]
number
[0111] Next, we use the tagged data (corresponding to training data) to calculate each α i Update as shown in (1) and (2) below. This update is repeated for each tagged data (corresponding to the training data). This single operation is counted as one epoch.
[0112] (1) Updating each value: Update according to the following formula 5. Here, ε is the learning rate (hyperparameter). This learning rate determines how much weight (weight parameter α i ) is modified.
[0113]
number
[0114] (2) Normalization: Normalize according to the following formula 6.
[0115]
number
[0116] The above steps (1) and (2) are repeated M times. Here, M is the number of epochs. This allows each α i is obtained.
[0117] 12 shows an example of the weight value calculated for each pipeline (first generation unit 521 to Nth generation unit 524 → feature extraction unit 530). The recording control unit 54 (see FIG. 9) associates the weight value with the pipeline and stores them in the weight DB 600 (see FIG. 10).
[0118] In the above, the loss function L defined in the above formula 3 is the tag value vector t (t1, t2, ..., t 15 ) and the weighted expectation y(vector y(y1,y2,…,y 15 ) is shown as an example, but the definition of the error function is not limited to this and other definitions may be used. For example, the error function may be defined as binary cross entropy, mean squared error, mean absolute error, root-mean-square error (RMSE or RMSD), mean squared logarithmic error (MSLE), Huber loss, Poisson loss, hinge loss, Kullback-Leibler divergence, etc.
[0119] In addition, in the above, the weight of the weighted average α i As the initial value of , the value determined based on the above-mentioned Equation 2 is used. However, the weight of the weighted average α i For example, the non-uniform weighting α obtained by matching the tag information FQ1 to FQ4 generated by the feature extraction unit 530 with the tag information TD3 constituting the tagged data TD1 may be set to another value as the initial value of α. imay be used as the initial value. For example, the initial value dependency weakens with each increase in iteration, but if the initial value is set to a value close to the goal value, the goal can be reached more quickly. Therefore, by setting the initial value to a value close to the goal value, it becomes possible to execute the calculation process quickly. This makes it possible to reduce the amount of calculation. In FIG. 13, the weight α i Here is an example of setting the initial value of to a value close to the goal value.
[0120] FIG. 13 shows the weight α of the weighted average in the weight calculation process by the information processing device 500. i 13 is a diagram showing a schematic example of a calculation when setting an initial value of
[0121] As shown in FIG. 13(A), the weight calculation unit 540 calculates the difference value of the feature for each tag of the tag information FQ1 based on the text information TX1 generated by the first generation unit 521 and the tag information TD3. For example, the feature value "0.4" of the tag "casual" in the tag information FQ1 is compared with the feature value "0.3" of the item "casual" in the tag information TD3, and "0.1" is calculated as the difference value. The difference values are calculated similarly for the other tags. Then, the weight calculation unit 540 calculates a total value by adding up the absolute values of the difference values of the feature calculated for each tag of the tag information FQ1 and the tag information TD3. FIG. 13(A) shows an example in which "0.7" is calculated as the total value.
[0122] As shown in FIG. 13B, the weight calculation unit 540 calculates the difference value of the feature for each tag between the tag information FQ2 based on the text information TX2 generated by the second generation unit 522 and the tag information TD3. As in the example shown in FIG. 13A, the weight calculation unit 540 calculates the difference value of the feature for each tag between the tag information FQ2 and the tag information TD3, and calculates the sum of the absolute values of the difference values. FIG. 13B shows an example in which the sum is calculated as "1.2". Although not shown, the difference value of the feature from the tag information TD3 is calculated for the tag information FQ3 and FQ4 based on the text information TX3 and TX4 generated by the other generation units (the third generation unit 523 and the Nth generation unit 524) in the same way, and the sum of the absolute values of the difference values is calculated.
[0123] 13C shows the sum of the absolute values of the difference values calculated by the above-mentioned calculation processes. In this case, the weight calculation unit 540 calculates the weight α i For example, it is possible to determine the initial value of the weight α i It is possible to set the initial value of each weight α to a small value, as in the above-mentioned formula 2. i Each weight α i The initial value of the weight α i The calculation process after setting the initial value is the same as the calculation process described above, so the explanation will be omitted here. In addition, when there are multiple tagged data (corresponding to teacher data), it is possible to use values obtained by performing a predetermined calculation process (e.g., averaging process) on each difference value calculated for each tagged data.
[0124] [Example of operation of information processing device] Fig. 14 is a flowchart showing an example of weight calculation processing in the information processing device 500. This weight calculation processing is executed based on a program stored in the storage unit 55. This weight calculation processing is executed when tagged data TD1 is input to the acquisition unit 510 by a user operation. This weight calculation processing will be described with appropriate reference to Figs. 9 to 13.
[0125] In step S701, the acquisition unit 510 accepts tagged data TD1 input by user U1, outputs image information TD2 constituting the accepted tagged data TD1 to the text information generation unit 520, and outputs tag information TD3 constituting the tagged data TD1 to the weight calculation unit 540.
[0126] In step S702, each of the generating sections (first generating section 521 to Nth generating section 524) constituting text information generating section 520 generates text information TX1 to TX4 (see FIG. 11) based on image information TD2 output from acquisition section 510.
[0127] In step S703, the feature extraction unit 530 generates tag information FQ1 to FQ4 (see FIG. 11) based on the text information TX1 to TX4 generated by the text information generation unit 520.
[0128] In step S704, the weight calculation unit 540 performs calculation processing of the weight value (weight value for each pipeline) for each generation unit (the first generation unit 521 to the Nth generation unit 524) based on the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the tag information TD3 constituting the tagged data TD1. This calculation processing is similar to the calculation processing described above, so a description thereof will be omitted here.
[0129] In step S705, the recording control unit 54 associates the weight calculated for each piece of tag information FQ1 to FQ4 (for each pipeline) in step S704 with each generation unit (first generation unit 521 to Nth generation unit 524) constituting the text information generation unit 520, and stores them in the weight DB 600. For example, each piece of information is stored in the weight DB 600 as shown in FIG.
[0130] [Example of generating tag information using weights] Next, an example of generating tag information using the above-mentioned weights will be described. Figures 15 and 16 show an example of generating tag information by executing a weighted average calculation process.
[0131] [Example of configuration of information processing device] Fig. 15 is a block diagram showing an example of a functional configuration of an information processing device 560. The information processing device 560 is a partial modification of the information processing device 500 shown in Fig. 9, and is common to the information processing device 500 except that an image acquisition unit 51 and a weighted average calculation unit 550 are provided instead of the acquisition unit 510 and the weight calculation unit 540, and an image DB 200 is added. For this reason, the parts common to the information processing device 500 are given the same reference numerals as those of the information processing device 500, and some of the descriptions thereof will be omitted.
[0132] The information processing device 560 includes an image acquisition unit 51, a text information generation unit 520, a feature extraction unit 530, a weighted average calculation unit 550, a recording control unit 54, and a storage unit 55. Each of the image acquisition unit 51, the text information generation unit 520, the feature extraction unit 530, the weighted average calculation unit 550, and the recording control unit 54 is realized by, for example, one or more processing circuits such as a CPU, a GPU, etc. Also, like the example shown in Fig. 2, the information processing device 10 and the information processing device 560 are shown as separate entities, but they may be configured as an integrated device.
[0133] The weighted average calculation unit 550 executes a weighted average calculation process based on the feature amount for each tag related to each of the multiple pieces of text information output from the feature extraction unit 530 and the weight value stored in the weight DB 600. Then, the weighted average calculation unit 550 outputs the feature amount for each tag related to the image (image acquired by the image acquisition unit 51) calculated by the weighted average calculation process to the recording control unit 54. Specifically, the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the weight α stored in the weight DB 600 are used to calculate the weighted average. i For example, for the tag "casual", weights α i The sum of the values multiplied by α is set as the feature value of the tag “casual”. Similarly, for other tags, the feature values of the tag information FQ1 to FQ4 and the weights α i The sum of the values obtained by multiplying each of these is regarded as the feature value of the tag.
[0134] [Example of operation of information processing device] Fig. 16 is a flow chart showing an example of feature extraction processing in the information processing device 560. This feature extraction processing is executed based on a program stored in the storage unit 55. This feature extraction processing shows an example of feature extraction processing using a weighted average. This feature extraction processing is executed when an image is acquired by the image acquisition unit 51. This weight calculation processing will be described with reference to Figs. 1 to 15 as appropriate.
[0135] In step S711, the image acquisition unit 51 acquires an image input by the user U1, and outputs to the text information generation unit 520 image information relating to the acquired image.
[0136] In step S712, each generating unit (first generating unit 521 to Nth generating unit 524) constituting text information generating unit 520 generates a plurality of pieces of text information based on the image information output from image acquiring unit 51, and outputs the generated plurality of pieces of text information to feature extracting unit 530. The method of generating the plurality of pieces of text information is similar to the example shown in FIG.
[0137] In step S713, feature extraction unit 530 extracts feature amounts for each of the multiple tags stored in tag DB 100, for each piece of text information, based on each of the multiple pieces of text information generated by text information generation unit 520. Then, feature extraction unit 530 outputs the feature amounts for each of the multiple tags extracted for each piece of text information to weighted average calculation unit 550. The method of extracting feature amounts for each tag is similar to the example shown in FIG.
[0138] In step S714, the weighted average calculation unit 550 performs weighted average calculation processing on the feature amounts for each of the multiple tags extracted for each piece of text information by the feature extraction unit 530 to obtain a set of tag information (feature amounts for each of the multiple tags). Then, the weighted average calculation unit 550 outputs the obtained feature amounts for each of the multiple tags to the recording control unit 54.
[0139] In step S715, the recording control unit 54 stores the feature amounts for each of the multiple tags output from the weighted average calculation unit 550 and the image information output from the image acquisition unit 51 in the image DB 200. For example, as shown in Fig. 4, the feature amounts for each of the multiple tags and the image information are stored in the image DB 200 in association with each other.
[0140] [Example of effect of the second embodiment] As described above, when one text information generating unit model (e.g., img2txt) is used, it is highly dependent on the features of the model, whereas using multiple text information generating unit models (e.g., img2txt) makes it possible to alleviate this dependency. In addition, using multiple models makes it possible to stabilize behavior. In addition, for example, in order to generate a trained model for extracting image features from a specific viewpoint according to a user's preferences, a large amount of training data with features associated by human work is required. In contrast, in the second embodiment, it is possible to calculate weights using only one or a small amount of tagged data (corresponding to training data). That is, in the second embodiment, it is possible to extract appropriate features even with a small amount of data.
[0141] [Example of correcting text information using weights] In the above, an example has been shown in which a weighted average calculation is performed on the feature amount for each tag based on the multiple text information generated by the multiple first generating units 521 to Nth generating units 524 using weights calculated using one or a small amount of tagged data (corresponding to teacher data). However, the feature amount for each tag may be calculated using these weights in other ways. For example, weights may be used for the multiple text information generated by each of the multiple first generating units 521 to Nth generating units 524. For example, it is assumed that the text information generated by a generating unit with a large weight has a relatively high accuracy, and the text information generated by a generating unit with a small weight has a relatively low accuracy. Therefore, for example, the text information generated by a predetermined number (e.g., 1 to 3) of generating units with a large weight is used as reference text information, and instruction information to correct one or more reference text information using text information generated by other generating units and to extract feature amounts for each tag for one or more text information after correction may be output to the feature extracting unit 530, and feature amounts for each tag may be extracted for one or more text information after correction. In this way, it is possible to generate corrected text information with higher accuracy by correcting reference text information that is estimated to be highly accurate with other text information. It is possible to extract feature amounts for each tag with higher accuracy by extracting feature amounts for each tag based on this corrected text information.
[0142] In other words, for example, for each of the pieces of text information TX1 to TX4 (see FIG. 11) generated by the generation units (first generation unit 521 to Nth generation unit 524) constituting the text information generation unit 520, new text information may be generated taking into consideration the weights, and features for each tag may be extracted based on the new text information.
[0143] For example, the importance of each sentence is set according to the weight for the first sentence included in the text information TX1, the second sentence included in the text information TX2, the third sentence included in the text information TX3, and the fourth sentence included in the text information TX4. For example, assume that the weight of the first generation unit 521 is the highest, the weight of the second generation unit 522 is the second highest, the weight of the third generation unit 523 is the third highest, and the weight of the Nth generation unit 524 is the lowest. In this case, the feature extraction unit 530 can generate a new sentence with reference to the second to fourth sentences, using the first sentence as a reference. Therefore, for example, it is possible to output the text information TX1 to TX4 to the feature extraction unit 530, and output instruction information to the feature extraction unit 530 instructing the unit 530 to generate a new sentence with reference to the second to fourth sentences, using the first sentence as a reference, and extract a feature amount from the new sentence.
[0144] As a result, it is possible to omit the weighted average calculation unit 550 and generate tag information taking the weights into consideration by using two or more types of text information generation units 520 and one type of feature extraction unit 530. When there are multiple pieces of corrected text information, it is possible to perform a weighted average calculation process using the weights to generate feature amounts for each set of tags.
[0145] [Example of setting tags based on text information] In the first embodiment, an example of setting a plurality of tags based on text information generated from an image is shown. Here, an example of setting a plurality of tags based on at least one of a plurality of text information generated by a plurality of first generating units 521 to Nth generating units 524 is shown. As described above, it is assumed that text information generated by a generating unit with a large weight has a relatively high accuracy, and text information generated by a generating unit with a small weight has a relatively low accuracy. Therefore, for example, one or a plurality of pieces of text information generated by a predetermined number (for example, 1 to 3) of generating units with a large weight may be used to set tags based on character strings, words, etc. contained in the text information. Note that the setting method of setting tags using text information is similar to the example shown in the first embodiment (an example of setting a plurality of tags based on text information generated from an image), and therefore a description thereof will be omitted here.
[0146] [Variations] In the second embodiment, an example is shown in which tag information is generated using two or more types of text information generating units 520 and one type of feature extracting unit 530. However, tag information may be generated using both two or more types of text information generating units and two or more types of feature extracting units. In this case, it is possible to calculate weights for each combination of the text information generating unit and the feature extracting unit, and generate tag information using the weights.
[0147] In the above, an example has been shown in which tag information is generated in consideration of weights using two or more types of text information generation units 520, one type of feature extraction unit 53, and weighted average calculation unit 550. However, the weighted average calculation unit 550 may be omitted, and tag information may be generated using two or more types of text information generation units 520 and one type of feature extraction unit 53. For example, text information combining each sentence included in each piece of text information TX1 to TX4 (see FIG. 11) generated by the generation units (first generation unit 521 to Nth generation unit 524) constituting the text information generation unit 520 may be output to the feature extraction unit 530.
[0148] For example, it is possible to output a combined sentence in which the first sentence to the fourth sentence are simply arranged using a first sentence included in the text information TX1, a second sentence included in the text information TX2, a third sentence included in the text information TX3, and a fourth sentence included in the text information TX4 to the feature extraction unit 530. In this case, the feature extraction unit 530 can extract a feature based on the combined sentence in which the first sentence to the fourth sentence are arranged. Also, for example, the combined sentence may be output to the feature extraction unit 530, and instruction information for executing a predetermined process on the combined sentence may be output to the feature extraction unit 530. For example, it is possible to extract a common part (e.g., a character string such as a word or a sentence) from the combined sentence (first combined sentence), combine the extracted part with a non-common part (e.g., a character string such as a word or a sentence) to generate a new combined sentence (second combined sentence), and output instruction information for instructing to extract a feature from the second combined sentence. Also, for example, it is possible to generate a new sentence based on the combined sentence (first combined sentence) in accordance with some criteria, and output instruction information instructing to extract features from the new sentence (fifth sentence).
[0149] [Variations] 1 and 2 show an example in which the information processing device 10 that generates a tag and the information processing device 50 that generates tag information using the tag are different devices. However, as described above, the device that generates a tag and the device that generates tag information using the tag may be the same device. In addition, in the above, an example in which the generated tag information is associated with an image and stored in the image DB 200 has been shown, but the generated tag information may be displayed together with the image, or may be associated with the image and transmitted to another device. Therefore, in FIGS. 17 and 18, an example is shown in which a processing unit that generates a tag and a processing unit that generates tag information using the tag are provided in the same device. In addition, in FIGS. 17 and 18, an example is shown in which the generated tag information is displayed together with an image.
[0150] [Example of configuration of information processing device] FIG. 17 is a block diagram showing an example of a functional configuration of the information processing device 800.
[0151] The information processing device 800 is an example of an information processing device having the configuration of the information processing device 10 shown in Fig. 1 and the configuration of the information processing device 50 shown in Fig. 2. Note that parts common to the information processing device 10 shown in Fig. 1 and the information processing device 50 shown in Fig. 2 are given the same reference numerals as the information processing devices 10 and 50, and some of the descriptions thereof will be omitted.
[0152] The information processing device 800 includes a DB control unit 810 and an output unit 820. The DB control unit 810 executes recording control and readout control for each DB in the storage unit 55. For example, the DB control unit 810 stores a tag list output from the tag setting unit 12 in the tag DB 100. The DB control unit 810 also associates the feature amount for each tag output from the feature extraction unit 53 with the image output from the image acquisition unit 51 and stores them in the image DB 200. The DB control unit 810 also associates the image information and tag information stored in the image DB 200 with each other and outputs them from the output unit 820. An example of this output will be described in detail with reference to FIG. 18.
[0153] The output unit 820 is an output unit that outputs various information (for example, a display unit 821 (see FIG. 18) and an audio output unit (not shown)). Specifically, the output unit 820 displays various information on the display unit 821 and outputs audio information from the audio output unit based on information output from the DB control unit 810. An image display device such as a display capable of outputting images and audio may be used as the output unit 820. Also, for example, an output device capable of outputting at least one of images and audio may be used. Note that, in FIG. 17, an example in which the display unit 821 and the audio output unit are provided in the information processing device 800 as an example of the output unit 820 is shown, but an output device separate from the information processing device 800 may be used as the output unit 820.
[0154] [Example of tag information display] Fig. 18 is a diagram showing a display example in which tag information generated by information processing device 800 is displayed on display unit 821. As shown in Fig. 18, it is possible to associate an image MG1 acquired by image acquisition unit 51 with a feature amount (tag information TG10) extracted by feature extraction unit 53 and display them on display unit 821. This allows user U1 to easily check tag information TG10 generated for a desired image MG1.
[0155] [Example of communication system configuration] In the above, an example has been shown in which a feature amount for each tag is extracted using a device that the user U1 can access. However, various input operations, etc. may be performed in a first device, and various processes corresponding to the input operations, etc. (e.g., feature extraction process for each tag) may be performed in a second device. In addition, various processes according to this embodiment may be performed using three or more devices. Thus, in Figs. 19 and 20, an example of a communication system TS1 capable of exchanging various information using multiple devices that can be connected via a network NW1 is shown.
[0156] FIG. 19 is a block diagram showing an example of a functional configuration of the communication system TS1.
[0157] The communication system TS1 includes a network NW1, information processing devices 900 and 920, an electronic device 930, and the like. For example, the network NW1, the information processing devices 900 and 920, the electronic device 930, and the like are connected to each other via the network NW1. Note that communication between these devices is performed using wired communication or wireless communication. Note that communication between these devices may be performed directly between the devices in addition to communication via the network NW1.
[0158] The network NW1 is a network such as a public line network, the Internet, etc. Furthermore, each device constituting the communication system TS1 is connected to the network NW1 by a communication method using wireless communication or a communication method using wired communication, or by both methods.
[0159] The information processing device 900 corresponds to the information processing device 800 shown in Fig. 17. Moreover, each unit in the information processing device 900 corresponds to each unit of the same famous place shown in Fig. 17. The information processing device 900 can be, for example, a server capable of providing various information.
[0160] The communication unit 910 exchanges various information with other devices by using at least one of wired communication and wireless communication. For example, the communication unit 910 executes a reception process for receiving various information transmitted from the information processing device 920 and the electronic device 930, a transmission process for transmitting various information to the information processing device 920 and the electronic device 930, and the like.
[0161] The information processing device 920 and the electronic device 930 are fixed or portable information processing devices owned by the user U1, such as a smartphone, a tablet terminal, a smart watch, or a personal computer. The information processing device 920 and the electronic device 930 are devices capable of wired or wireless communication with the information processing device 900. The information processing device 920 and the electronic device 930 perform a transmission process for transmitting various information to the information processing device 900 and a reception process for receiving various information from the information processing device 900. For example, the information processing device 920 and the electronic device 930 can display various images on the display units 921 and 931 and output various sounds based on the information received from the information processing device 900.
[0162] FIG. 20 shows an example of transition of a display screen displayed on the display unit 931 of the electronic device 930. In FIG.
[0163] FIG. 20(A) shows an example of a display screen 932 displayed when setting a tag and an image to be tagged. The display screen 932 displays an input field display area 933, a decision button 934, and an image list display area MG10. The information input into the input field display area 933 is the same as the information input into the character input field 17 shown in FIG. 5(A). The images displayed and selected in the image list display area MG10 are the same as the image group 40 shown in FIG. 2 and the like. For example, when a character string "clothes" desired by the user U1 is input into the input field display area 933, and the image desired by the user U1 is selected from among the images displayed in the image list display area MG10, and then a selection operation is performed to select the decision button 934, character information related to the character string "clothes" and image information related to the selected image are transmitted to the information processing device 900. When the information processing device 900 receives each piece of information, the information processing device 900 generates tag information related to the selected image based on each piece of information. The method for generating this tag information is the same as the above-mentioned generation method. Then, the information processing device 900 executes display control to display the generated tag information and image information related to the selected image on the electronic device 930. An example of the display in this case is shown in FIG. 20(B).
[0164] Fig. 20(B) shows an example of a display screen 935 on which an image MG1 and tag information TG10 are displayed in association with each other. Note that the image MG1 and tag information TG10 are similar to the example shown in Fig. 18. In this way, the tag information of an image can be easily obtained using a portable device or the like carried by user U1.
[0165] 17 and 19 show configuration examples in which the configuration shown in FIG. 1 and the configuration shown in FIG. 2 are combined, but a configuration example in which the configuration shown in FIG. 9 and the configuration shown in FIG. 15 are combined may also be used.
[0166] [Example of application to video] In the above, an example of extracting the feature amounts of a plurality of tags from a still image has been shown. However, this embodiment can also be applied to a video composed of a plurality of images (frames). For example, among the plurality of images (frames) constituting a video, the feature amounts of a plurality of tags are extracted from images at a predetermined interval (for example, images at intervals of several seconds to several tens of seconds), and a predetermined calculation (for example, calculation of an average value for each tag) is performed on the feature amounts of the extracted plurality of tags, thereby making it possible to extract the feature amounts of a plurality of tags related to the entire video. In this case, for example, an image (for example, a representative image) showing the video and the extracted tag information (feature amounts of a plurality of tags) may be associated and stored in the video DB, and they (the image showing the video, the extracted tag information) may be associated and displayed on the display unit, or they may be associated and transmitted to another device. This makes it possible to easily grasp the tag information related to one video. Also, for example, for a video including a plurality of scenes (for example, a video in which the shooting scene is changed sequentially), the above-mentioned predetermined calculation process may be performed for each scene, and the feature amounts of a plurality of tags related to each scene may be extracted. In this case, for example, images showing each scene (e.g., a representative image for each scene) and tag information extracted for each scene (feature amounts of multiple tags) may be associated and stored in the video DB, and these (images showing each scene, tag information extracted for each scene) may be associated and displayed on a display unit, or may be associated and transmitted to another device. This makes it possible to easily grasp the tag information for each scene related to one video.
[0167] [Third embodiment] [Example of searching for the image the user desires] The above describes an example of tagging processing in which a plurality of tag feature quantities are extracted from an image and tagged. Images tagged in this manner can be used to search for images that are to the user's taste, images that the user desires, etc. Therefore, an example of a search process for images that are to the user's taste, images that the user desires, etc. will be described below.
[0168] [Example of information processing system configuration] FIG. 21 is a block diagram showing an example of the configuration of an information processing system IS1.
[0169] The information processing system IS1 is composed of an information processing device 1000 and an electronic device 1100. For example, the information processing device 1000 and the electronic device 1100 are connected via a predetermined network NW1. Note that communication between these devices is performed using wired communication or wireless communication. Furthermore, communication between these devices may be performed directly between the devices in addition to communication via the network NW1.
[0170] [Server configuration example] The information processing device 1000 includes a communication unit 1010, a control unit 1020, and a storage unit 1030. The information processing device 1000 can be realized by an information processing device or electronic device such as a server, a personal computer, a smartphone, or a tablet terminal. The control unit 1020 is realized by, for example, one or more processing circuits such as a CPU, a GPU, or the like.
[0171] The communication unit 1010 exchanges various information with other devices by using at least one of wired communication and wireless communication. For example, the communication unit 1010 executes a reception process for receiving various information transmitted from the electronic device 1100, a transmission process for transmitting various information to the electronic device 1100, and the like.
[0172] The control unit 1020 includes an information acquisition unit 1021 , a tag generation unit 1022 , a vector generation unit 1023 , an extraction unit 1024 , and an information provision unit 1025 .
[0173] The information acquisition unit 1021 acquires various pieces of information transmitted from the electronic device 1100, and supplies the acquired pieces of information to each unit.
[0174] The tag generating unit 1022 generates a tag for the information (e.g., image information, text information) output from the information acquiring unit 1021, and outputs information on the generated tag (tag information) to the vector generating unit 1023. For example, when the image information output from the information acquiring unit 1021 is an image stored in the image information DB 1040, the tag generating unit 1022 acquires a feature amount 1044 (see FIG. 23) associated with the image. The identity of the images can be determined based on the identification information associated with the images.
[0175] Furthermore, when the information output from the information acquisition unit 1021 is other image information (an image not stored in the image information DB 1040) or text information, the tag generation unit 1022 can generate a tag based on tag information stored in the tag information DB 1060. Furthermore, when the information output from the information acquisition unit 1021 is other image information or text information, the tag generation unit 1022 can generate a tag using an AI (Artificial Intelligent) model. The tag generation method will be described in detail with reference to Figs. 25 to 29 and the like.
[0176] The vector generating unit 1023 generates a user vector according to the taste of the user of the electronic device 1100 based on the tag information (vector components) output from the tag generating unit 1022, and outputs the generated user vector to the extracting unit 1024. The method of generating the user vector will be described in detail with reference to Figs. 25 to 29 and the like.
[0177] The extraction unit 1024 extracts desired information from each DB stored in the storage unit 1030 based on the information output from each unit, and outputs each extracted information to the information providing unit 1025. For example, when a user vector is output from the vector generating unit 1023, the extraction unit 1024 compares the user vector with a provider vector (tag 1053, feature amount 1054 (see FIG. 24)) stored in the provider information DB 1050, and extracts provider information according to the preference of the user of the electronic device 1100 based on the similarity of each vector. Also, for example, when the electronic device 1100 requests the provision of a new image, the extraction unit 1024 compares the user vector generated by the vector generating unit 1023 with the feature amount 1044 (vector components) stored in the image information DB 1040, and extracts an image according to the preference of the user of the electronic device 1100 based on the similarity of each vector. The method of generating each of these user vectors will be described in detail with reference to FIG. 25 to FIG. 29 and the like.
[0178] The information providing unit 1025 executes transmission control for transmitting information output from each unit to the electronic device 1100 .
[0179] The storage unit 1030 is a storage medium that stores various types of information. For example, the storage unit 1030 stores various types of information (for example, a control program, image information 1040 (see FIG. 23), a provider information DB 1050 (see FIG. 24), a tag information DB 1060, and an AI model) required for the control unit 1020 to perform various processes. The storage unit 1030 also stores various types of information acquired via the communication unit 1010. For example, a ROM, a RAM, an SRAM, a HDD, an SSD, or a combination of these can be used as the storage unit 1030.
[0180] The image information DB 1040 is a database for storing images provided by providers in association with feature quantities of multiple tags generated based on the images. Here, the provider refers to, for example, an individual, a company, or the like that provides various products, various services, and the like using a website (for example, an EC (Electronic Commerce) site). The image information DB 1040 will be described in detail with reference to FIG. 23.
[0181] The provider information DB 1050 is a database for storing, in association with one another, various pieces of information about providers who provide various products, various services, and the like to users of the electronic devices 1100. The provider information DB 1050 will be described in detail with reference to FIG.
[0182] The tag information DB 1060 is a database for storing information used when tag information is generated by the tag generating unit 1022. For example, when tag information is generated using a dictionary database that associates various information including a character string, a tag corresponding to the character string, and a feature corresponding to the tag, information about the various information is stored. Also, for example, when tag information is generated using an AI model, information about the AI model is stored.
[0183] [Examples of electronic device configurations] The electronic device 1100 includes a communication unit 1110, a position information acquisition unit 1120, an image acquisition unit 1130, a sound acquisition unit 1140, a control unit 1150, a UI (User Interface) unit 1160, and a storage unit 1170. The electronic device 1100 is realized by, for example, an information processing device such as a smartphone, a tablet terminal, or a personal computer, an electronic device, or other device. Note that each unit included in the electronic device 1100 is an example, and some of these may be omitted, or other members may be added. For example, when the electronic device 1100 is a personal computer, the position information acquisition unit 1120 and the image acquisition unit 1130 can be omitted or can be external devices.
[0184] The communication unit 1110, under the control of the control unit 1150, exchanges various types of information with other devices using wired or wireless communication.
[0185] The location information acquisition unit 1120 acquires location information on the location where the electronic device 1100 exists, and outputs the acquired location information to the control unit 1150. The location information acquisition unit 1120 can be realized by, for example, a GPS receiver that receives a GPS (Global Positioning System) signal and calculates location information based on the GPS signal. The calculated location information includes data on the position such as latitude, longitude, and altitude at the time of receiving the GPS signal. A location information acquisition device that acquires location information by other location information acquisition methods may be used. For example, a location information acquisition device that derives location information using access point information by wireless LAN present in the vicinity and acquires this location information may be used. For example, a location information acquisition device that derives location information using position information of a base station used for calls and communications may be used. For example, a location information acquisition device that derives location information using a position estimation technology by a navigation function and acquires this location information may be used. For example, the position of the device itself can be estimated based on sensor information from various sensors (for example, an acceleration sensor, a gyro sensor) and map information.
[0186] The image acquisition unit 1130 captures an image of a subject and generates an image (image data) based on the control of the control unit 1150, and outputs the generated image to the control unit 1150. The image acquisition unit 1130 is configured with, for example, an imaging element (image sensor) into which light from the subject collected by a lens (not shown) is incident, and an image processing unit that performs predetermined image processing on the image data generated by the imaging element. For example, a CCD (Charge Coupled Device) type or a CMOS (Complementary Metal Oxide Semiconductor) type imaging element can be used as the imaging element.
[0187] The sound acquisition unit 1140 acquires sounds around the electronic device 1100 under the control of the control unit 1150, and outputs sound information related to the acquired sounds to the control unit 1150. As the sound acquisition unit 1140, for example, one or more microphones or sound acquisition sensors can be used.
[0188] The control unit 1150 controls each unit based on various programs stored in the storage unit 1170. The control unit 1150 is realized by a processing device such as a CPU, a GPU, etc. Each process executed by the control unit 1150 will be described in detail with reference to Figs. 26 to 29 and the like.
[0189] The storage unit 1170 is a storage medium that stores various types of information. For example, the storage unit 1170 stores various types of information (for example, a control program, an image information DB 1080 (see FIG. 33), a search application, and device identification information) required for the control unit 1150 to perform various processes. The storage unit 1170 also stores various types of information acquired via the communication unit 1110. For example, a ROM, a RAM, an SRAM, a HDD, an SSD, or a combination of these can be used as the storage unit 1170. The storage unit 1170 may be a removable storage unit or may be a storage unit built into the electronic device 1100. The electronic device 1100 may be provided with both a removable storage unit and a built-in storage unit.
[0190] The UI unit 1160 includes a reception unit 1161 and an output unit 1162 .
[0191] The reception unit 1161 receives various operations from the user and outputs the received operation contents to the control unit 1150. The reception unit 1161 and the output unit 1162 may be configured as a touch panel that allows the user to input operations by touching or approaching the display surface with their finger, or may be configured as a separate user interface. When configured as a separate user interface, various operation members such as switches, buttons, and keyboards can be used as the reception unit 1161. Also, a user interface may be configured in which an image corresponding to the operation unit is projected onto a wall or the like and the image is used to perform operations (for example, the user points at it). Also, the sound acquisition unit 1140 or the like that receives various operations based on the voice uttered by the user may be used as the reception unit 1161.
[0192] The output unit 1162 outputs various information based on the control of the control unit 1150. For example, the output unit 1162 can be configured with a display unit and a sound output unit. This display unit displays various images based on an instruction from the control unit 1150. For example, a display panel such as an organic EL (Electro Luminescence) panel or an LCD (Liquid Crystal Display) panel can be used as the display unit. The sound output unit outputs various sounds based on an instruction from the control unit 1150. For example, one or more speakers can be used as the sound output unit. Note that the display unit, the sound output unit, the sound acquisition unit 1140, and the reception unit 1161 are examples of user interfaces, and some of them may be omitted, or other user interfaces may be used.
[0193] [Tagging example] Here, tagging applied to image information stored in the image information DB 1040 will be described. For example, image information uploaded to the information processing device 1000 by a provider using the service in this embodiment is stored in the image information DB 1040. That is, image information uploaded to the information processing device 1000 by the provider is registered as a provided image. When an image is registered in this manner, a tagging process is executed for the image. Then, the registered image and a tag generated for it are stored in the image information DB 1040 in association with each other.
[0194] Here, as described above, a tag is a criterion for classifying images based on the properties of each image, and also means an item that serves as a criterion for the classification. That is, tagging an image can be referred to as giving meaning to the image. A tag also means an item that serves as a criterion when extracting the features (feature amount) of an image. A tag can also be referred to as information for identifying an image. Note that a label described below also means information that represents the features, attributes, etc. of an image. In this way, a tag and a label have approximately the same or similar meanings, and therefore, hereinafter, a tag and a label may be used to mean the same thing.
[0195] In addition, in this embodiment, as described above, information in which a tag and a corresponding feature (feature amount) are associated with each other will be referred to as tag information. However, tag information can also be referred to as a tag dictionary, metadata, associated information, additional information, etc. In addition, in this embodiment, an example will be described in which a quantified feature amount is used as a feature corresponding to a tag.
[0196] For example, tags that can be added to an image of a mug (i.e., tags that represent the taste of the mug's design) can include simple, modern, retro, classic, casual, fancy, natural, art deco, bohemian, floral, animal motif, ethnic, Nordic, etc.
[0197] Each of these tags may be expressed, for example, by presence / absence (for example, 0 or 1), or may be expressed in multiple stages (for example, five stages from 0 to 4, a range from 0 to 1). FIG. 23 and other figures show an example in which each tag is expressed in multiple stages (a range from 0 to 1).
[0198] [Example of tagging done by humans] When an image is uploaded to the information processing device 1000 by a provider, it is possible to execute a tagging process in which a tag is manually assigned to the image. For example, a person can judge the elements of the object contained in the image (for example, the style and taste of the object contained in the image) and assign the above-mentioned tag information to the image based on the judgment result. However, when tags are assigned manually in this manner, personnel with specialized knowledge are required at the time of registration, and it is often difficult to obtain such personnel. Therefore, an example of executing the tagging process using artificial intelligence (AI) or the like is shown below. Note that in this embodiment, an example of assigning tags mainly using AI is described.
[0199] [Example of tagging using AI] [Examples of images and teacher labels used for learning] FIG. 22(A) shows an example of image IMG1 used for learning. Here, an image of a mug will be used as image IMG1. That is, when executing a tagging process that uses AI to tag images, various images such as image IMG1 can be used as training data. Note that image IMG1 is a simplified image for ease of explanation.
[0200] Note that a tag pattern may be prepared for each image category (e.g., mug, sofa furniture, house) and an AI model may be generated for each image category, or a common tag pattern may be prepared regardless of the image category and an AI model may be generated. Figure 22(B) shows an example of tags to be added as teacher labels to an image IMG1 used for learning.
[0201] 22B, when learning image IMG1, predetermined tags are assigned to image IMG1 as teacher labels, and learning is then performed using the assigned tags.
[0202] FIG. 22B also shows an example of using tags that are quantified between 0 and 1 to indicate the degree to which the tag applies. In this case, a tag value of 0 means that the content applies to the lowest extent, and a tag value of 1 means that the content applies to the highest extent. For example, in FIG. 22B, the value (feature amount F1) of tag TI1 "Art Deco" is set to "0.9". This means that image IMG1 (mug) has a very high degree of "Art Deco". In other words, the mug included in image IMG1 gives a strong impression of Art Deco style.
[0203] Similarly, for images other than the image IMG1 used for learning, predetermined tags are added as teacher labels and used for learning.
[0204] [Example of generating an AI model for learning] An AI model (e.g., a machine learning model generated by machine learning) is generated using the above-mentioned teacher data (e.g., image IMG1) and teacher labels (e.g., tag TI1, feature F1). For example, the teacher data (e.g., image IMG1) and teacher labels (e.g., tag TI1, feature F1) can be read and trained in advance, and this AI model can be used for tagging processing. Note that, for example, CNN (Convolutional Neural Network), ViT (Vision Transformer), etc. can be used as the AI model. Note that these can be executed by installing a predetermined application in the information processing device 1000.
[0205] [Example of output from a trained AI model] As described above, when a tag pattern is prepared for each image category (for example, mug, furniture sofa, house) and an AI model is generated for each image category, multiple AI models are prepared for each tag pattern. In this case, an image registered in the information processing device 1000 is input to an AI model corresponding to the image category. As a result, the AI model corresponding to the image category outputs a tagging result corresponding to the category. In this way, it is possible to use multiple AI models to output an input image with tag information related to the set tag. Also, as described above, when an AI model is generated by preparing a tag pattern regardless of the image category, one AI model is prepared. In this case, an image registered in the information processing device 1000 is input to the AI model. As a result, the AI model outputs a tagging result. In this way, it is possible to use one AI model to output an input image with tag information related to the set tag.
[0206] [Image information DB configuration example] Fig. 23 is a simplified diagram showing the contents of image information stored in image information DB 1040. Fig. 23 shows an example of storing image information of a mug as an example of an image category. Note that, for ease of explanation, Fig. 23 shows an example in which an image information DB is prepared for each image category (e.g., mug, sofa furniture, house) and image information is stored for each image category, but is not limited to this. For example, one or more common image information DBs may be prepared regardless of image category and various types of image information may be stored therein.
[0207] The image information DB 1040 is a database for storing images provided by providers in association with the feature amounts of a plurality of tags generated based on the images.
[0208] Specifically, image information 1042 , a tag 1043 , a feature amount 1044 , and provider identification information 1045 are stored in the image information DB 1040 in association with image identification information 1041 .
[0209] The image identification information 1041 is identification information for identifying each image prepared by the provider to be provided to the user. Note that, for ease of explanation, in Fig. 23, an example is shown in which only numerical values are stored as identification information in the image identification information 1041, but other identification information may be used.
[0210] The image information 1042 is information about an image prepared by a provider to provide to a user. In Fig. 23, for ease of explanation, an example in which only an image is stored in the image information 1042 is shown, but various attribute information stored in association with the image may also be included.
[0211] The tag 1043 is information indicating a tag given to an image stored in the image information 1042. For example, if the image stored in the image information 1042 is an image of a mug, as described above, simple, modern, retro, classic, casual, fancy, natural, art deco, bohemian, floral, animal motif, ethnic, Nordic, etc. are stored.
[0212] The feature amount 1044 is information indicating the feature amount assigned to each tag for the image stored in the image information 1042. Specifically, tag information assigned by the above-mentioned tag assignment process is stored.
[0213] For example, when images stored in the image information DB 1040 are displayed on the UI unit 1160 (see FIG. 21) of the electronic device 1100, the feature amount 1044 can be used to narrow down images that the user likes. For example, if a user is particular about fancy designs, the UI unit 1160 can display search results limited to images with a high value of "fancy" based on the feature amount 1044 associated with the tag 1043 "fancy". In this case, the user can select a preferred image from among images displayed on the UI unit 1160 (for example, images with a value of a predetermined value or more given to the feature amount 1044 corresponding to the tag 1043 "fancy"). This makes it possible to generate a user vector (described later) that reflects the user's preferences more. Also, for example, if a user wants to be recommended only providers who deal in Art Deco-style mugs, the UI unit 1160 can display search results limited to images with a high value of "Art Deco" based on the feature amount 1044 associated with the tag 1043 "Art Deco". In this case, the user can select a preferred image from among the images displayed on the UI unit 1160 (images with a high "Art Deco" value). This makes it possible to more appropriately recommend providers who deal in Art Deco-style mugs. Note that illustrations and other explanations regarding the narrowing down of searches for these images, provider information, etc. are omitted.
[0214] The provider identification information 1045 is identification information for identifying the provider who provided the image stored in the image information 1042. Note that, for ease of explanation, Fig. 23 shows an example in which the provider's name is stored in the provider identification information 1045, but other information capable of identifying the provider (for example, a serial number, a code) may be used.
[0215] Note that the information shown in FIG. 23 is an example of information stored in the image information DB 1040, and other information may be stored in association with the information, and some of the information may be omitted as necessary.
[0216] [Example of provider information DB configuration] FIG. 24 is a simplified diagram showing the contents of the provider information stored in the provider information DB 1050. As shown in FIG.
[0217] The provider information DB 1050 is a database for storing various information related to providers in association with each other.
[0218] Specifically, provider identification information 1051, profile information 1052, tag 1053, and feature amount 1054 are stored in the provider information DB 1050 in association with each other.
[0219] The donor identification information 1051 is identification information for identifying a donor. The donor identification information 1051 corresponds to the donor identification information 1045 shown in Fig. 23. Similarly to the example shown in Fig. 23, in order to facilitate explanation, Fig. 24 shows an example in which the name of the donor is stored in the donor identification information 1051, but other information capable of identifying the donor (for example, a serial number, a code) may be used.
[0220] The profile information 1052 is various information related to the provider whose provider identification information is stored in the provider identification information 1051. For example, the provider's address, provider introduction information, etc. are stored. This profile information can be edited by the provider as appropriate.
[0221] The tag 1053 is information (tag information) indicating a tag given to a donor whose donor identification information is stored in the donor identification information 1051. As this tag information, information regarding all or a part of the tag 1043 shown in Fig. 23 is stored.
[0222] The feature amount 1054 is information indicating a feature amount given to a provider whose provider identification information is stored in the provider identification information 1051. For example, a feature amount related to a product or service provided by the provider (a feature amount related to an image including the product or service) can be used as this feature amount. Note that this feature amount may be appropriately set by the provider, or may be stored based on the calculation result of each tag information given to a plurality of images uploaded to the information processing device 1000 by the provider. For example, an average value of each feature amount related to each tag given to a plurality of images uploaded to the information processing device 1000 can be stored in the corresponding feature amount. This feature amount is used when selecting a provider to be recommended to a user.
[0223] Note that each piece of information shown in FIG. 24 is an example of information stored in the provider information DB 1050, and other information may be stored in association with the information, and some of the information may be omitted as necessary.
[0224] [Example of selecting your favorite image] 25 is a diagram showing an example of a search screen 1200 displayed on the UI unit 1160 of the electronic device 1100. The search screen 1200 is displayed on the output unit 1162 under the control of the control unit 1150.
[0225] The search screen 1200 displays a character input area 1201 and a decision button 1202. Specifically, the information providing unit 1025 of the information processing device 1000 transmits search screen information for displaying the search screen 1200 to the electronic device 1100 based on a request from the electronic device 1100. When the electronic device 1100 receives the search screen information, the control unit 1150 causes the UI unit 1160 to display the search screen 1200 based on the received search screen information. Note that, in the character input area 1201, a user can input a desired character based on a user operation. FIG. 25 shows an example in which text information "I want a stylish mug" is input in the character input area 1201. In this way, when text information is input in the character input area 1201 and the decision button 1202 is selected, the control unit 1150 of the electronic device 1100 transmits search information including the text information to the information processing device 1000. When the information processing device 1000 receives the search information, the extraction unit 1024 executes a selection process to select an image to be provided to the electronic device 1100 from the image information DB 1040 based on the text information included in the received search information. For this selection process, a known selection technique for selecting an image from text information can be used. A selection process (see Figs. 36 to 38) described later may be used. Images in other devices (for example, an image providing server) may be searched for and used. An example of an image selected in this selection process is shown in Fig. 26. Note that, for ease of explanation, Fig. 25 shows an example in which image information DB 1040 prepared for each image category (for example, mug, furniture sofa, house) is used, but is not limited to this. For example, the selection process can be similarly applied to a case in which an image to be provided to the electronic device 1100 is selected using one or more image information DBs prepared regardless of the image category.
[0226] 26 is a diagram showing an example of a selection screen 1210 displayed on the UI unit 1160 of the electronic device 1100. The selection screen 1210 is displayed on the output unit 1162 under the control of the control unit 1150.
[0227] The selection screen 1210 displays an image display area 1211 that displays a plurality of images 1213 to 1220, and a decision button 1221. Specifically, the information providing unit 1025 of the information processing device 1000 transmits selection screen information for displaying an image selected by the selection process by the extraction unit 1024 to the electronic device 1100. When the electronic device 1100 receives the selection screen information, the control unit 1150 causes the UI unit 1160 to display the selection screen 1210 based on the received selection screen information. Note that other images can be displayed in the image display area 1211 based on a user operation (e.g., a scroll operation).
[0228] Furthermore, when a user operation is received by the receiving unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. For example, when a selection operation (e.g., a touch operation) for selecting a favorite image displayed in the image display area 1211 is performed by the user, the favorite image selected by the selection operation is set to a selected state. For example, the favorite image selected by the selection operation is displayed in a different manner from other images (e.g., by adding selection markers 1222 and 1223 (star regions), surrounding the image with a thick line, or adding a specific color) so that the user can understand the selected state. FIG. 26 shows an example in which the images 1218 and 1220 are selected by adding the selection markers 1222 and 1223. After the user performs the selection operation, only the image selected by the user may be displayed. It is also possible to cancel the selected state of the image by performing a selection operation (e.g., a click operation) on the selection markers 1222 and 1223 (star regions).
[0229] Furthermore, after one or more images are selected by the user, when the user executes a selection operation (for example, a touch operation) to select the decision button 1221, the control unit 1150 transmits image selection information related to the one or more images selected by the user to the information processing device 1000. When the information processing device 1000 receives the image selection information, the information acquisition unit 1021 outputs the received image selection information to the tag generation unit 1022 and the vector generation unit 1023.
[0230] [Example of generating user vectors] The tag generation unit 1022 acquires, from the image information DB 1040 (see FIG. 23 ), tag information associated with one or more images corresponding to the image selection information transmitted from the electronic device 1100. Specifically, the tag generation unit 1022 acquires, from the tag 1043 and the feature amount 1044 of the image information DB 1040, one or more pieces of tag information (feature amount) associated with the image identification information 1041 corresponding to the one or more images corresponding to the image selection information, and outputs the acquired tag information to the vector generation unit 1023.
[0231] Next, the vector generation unit 1023 generates a user vector related to the user who transmitted the image selection information, based on the tag information output from the tag generation unit 1022. Note that when there is one image corresponding to the image selection information transmitted from the electronic device 1100, the tag information associated with the one image can be used as the user vector.
[0232] For example, the vector generating unit 1023 can generate a user vector by performing an addition process on each element constituting each vector corresponding to the tag information output from the tag generating unit 1022 and dividing the addition result by the number of vectors to be added. Note that, in this embodiment, for ease of explanation, an example is shown in which a common tag (tag 1043 (see FIG. 23)) is set for each image regardless of the image category. Similarly, an example is shown in which a common tag (tag 1053 (see FIG. 24)) is set for each provider.
[0233] For example, assume that the user has selected image 3. The elements (components) constituting the tag information (tag vector) associated with the selected image 3 are [a1, ... , an1], [b1, ... , bn1], and [c1, ... , cn1]. Here, n1 means the number of items (components) of the tag. That is, it means the number of tags stored in the tag 1043 (see FIG. 23). For example, when simple, modern, retro, classic, casual, fancy, natural, art deco, bohemian, floral, animal motif, ethnic, and Nordic are used as tags, n1=13. Also, the components (numeric values) of each element are the values of the feature amount 1044 (see FIG. 23). For example, when a value in the range of 0 to 1 is used as the value of the feature amount 1044, the components (numeric values) of each element are values in the range of 0 to 1.
[0234] In this case, the vector generation unit 1023 generates a user vector by adding up the corresponding components of the three multidimensional vectors ([a1, ..., an1], [b1, ..., bn1], [c1, ..., cn1]) and dividing the sum by 3, resulting in a value [(a1+b1+c1) / 3, ..., (an1+bn1+cn1) / 3]. In this case, each element of the user vector corresponds to each element of the provider vector (each element of the feature amount 1054 (see FIG. 24)).
[0235] In this way, the average value of a plurality of images selected by the user can be used as the user vector. However, this embodiment is not limited to this, and the user vector can be generated by other calculation methods. For example, it is possible to adopt a calculation method in which an average is calculated after using a softmax function for each tag of each image. Also, for example, it is possible to adopt a calculation method in which an average is calculated using a Max function. Also, in this embodiment, an example is shown in which a user vector is generated using elements constituting each tag, but this is not limited to this, and the user vector may be generated using some elements of each tag, or the user vector may be generated using other tag information.
[0236] [Example of provider selection] Here, an example is shown in which an image provider is selected based on an image selected by a user. Based on the user vector generated by the vector generation unit 1023, the extraction unit 1024 extracts a provider according to the user's preference corresponding to the user vector from among the providers stored in the provider information DB 1050. For example, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with the provider vector stored in the feature amount 1054 of the provider information DB 1050, and extracts a provider according to the user's preference based on the comparison result.
[0237] Here, an example is shown in which a provider vector (feature 1054) is expressed as [S1, ..., Sn1]. Here, n1 means the number of items corresponding to each tag. Also, the components (numeric values) of each element correspond to the value of feature 1054.
[0238] In this case, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with a provider vector (feature amount 1054) [S1, ..., Sn1], and determines the similarity between them based on the comparison result. For example, if the user vector is [(a1+b1+c1) / 3, ..., (an1+bn1+cn1) / 3], the extraction unit 1024 compares this user vector with a provider vector [S1, ..., Sn1].
[0239] For example, the extraction unit 1024 calculates the cosine similarity between the user vector and each provider vector for each provider vector. A known calculation method can be used to calculate the cosine similarity. When calculating the cosine similarity, it is preferable to normalize each vector. Then, the extraction unit 1024 extracts the provider corresponding to the provider vector with the highest cosine similarity (i.e., the smallest cosine value with the user vector) as the provider according to the user's preference. Note that high cosine similarity means that the distance is short in vector space. In addition, when multiple providers are provided to the user, the extraction unit 1024 can extract a predetermined number (for example, about 2 to 10) of providers with high calculated cosine similarity. The number of providers extracted in this manner may be specified based on a user operation using the electronic device 1100.
[0240] Alternatively, the provider extraction process may be performed using other calculation methods. For example, the extraction unit 1024 calculates the difference (for each corresponding component) between each component constituting the user vector and each component constituting each provider vector (component corresponding to each component of the user vector) for each provider vector, and calculates the sum of these difference values for each provider vector. Then, the extraction unit 1024 may extract the provider corresponding to the provider vector with the smallest sum of the difference values as the provider according to the user's preference.
[0241] In this way, one provider corresponding to the provider vector most similar to the user vector may be selected, or multiple providers corresponding to a predetermined number of provider vectors having high similarity to the user vector may be selected in descending order of similarity. In other words, one provider corresponding to the provider vector closest to the user vector may be selected, or multiple providers corresponding to provider vectors having close distances may be selected in descending order.
[0242] Moreover, the extraction unit 1024 outputs the extracted information about the provider (eg, the profile information 1052, the image information 1042) to the information providing unit 1025.
[0243] Then, the information providing unit 1025 transmits to the electronic device 1100 selected provider information (for example, the profile information 1052 and the image information 1042) relating to the extracted provider.
[0244] [Example of selected provider information display] 27 is a diagram showing an example of a provider information screen 1230 displayed on the UI section 1160 of the electronic device 1100. The provider information screen 1230 is displayed on the UI section 1160 under the control of the control section 1150.
[0245] The provider information screen 1230 displays information about the provider extracted by the extraction unit 1024 of the information processing device 1000 (for example, profile information 1052, image information 1042). For example, the provider information screen 1230 displays a profile information display area 1231, an image display area 1232 displaying a plurality of images 1235 to 1237 associated with the provider extracted by the extraction unit 1024 of the information processing device 1000, a re-suggest new information button 1233, and a decision button 1234. Specifically, the information providing unit 1025 of the information processing device 1000 transmits screen information for displaying the provider information screen 1230 to the electronic device 1100. When the electronic device 1100 receives the screen information, the control unit 1150 causes the UI unit 1160 to display the provider information screen 1230 based on the received screen information. In addition, the image display area 1232 can display other images associated with the provider based on a user operation (for example, a scroll operation).
[0246] Furthermore, when a user operation is accepted by the accepting unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. For example, after checking the profile displayed in the profile information display area 1231 and the multiple images 1235 to 1237 displayed in the image display area 1232, when deciding on the provider of the profile displayed in the profile information display area 1231 as preferred information, the user performs a selection operation (e.g., a touch operation) to select the decision button 1234. When the selection operation of the decision button 1234 is performed, the control unit 1150 transmits decision information indicating that the selection operation has been performed to the information processing device 1000. When the information processing device 1000 receives the decision information, the control unit 1020 executes a predetermined process for enabling some kind of exchange between the provider corresponding to the received decision information and the user. As this predetermined process, for example, at least one of a transmission process of transmitting information about the user (information with the user's consent) to the provider corresponding to the determined information and a transmission process of transmitting information for accessing the provider corresponding to the determined information (for example, access information such as a telephone number, an email address, a URL (Uniform Resource Locator), or an SNS (Social Networking Service)) to the user is executed. This allows the user to easily and quickly access information according to his / her preferences.
[0247] Here, it is assumed that the user may request a re-proposal of other information after checking the profile displayed in the profile information display area 1231 and the multiple images 1235 to 1237 displayed in the image display area 1232. For example, it is conceivable that the user, having checked the multiple images 1235 to 1237 displayed in the image display area 1232, may request a re-proposal because they do not match his / her preferences.
[0248] In this way, when the user requests re-proposal of other information, the user performs a selection operation (for example, a touch operation) to select the re-proposal button 1233 for new information. When the selection operation of the re-proposal button 1233 for new information is performed, the control unit 1150 transmits re-proposal request information indicating that the selection operation has been performed to the information processing device 1000. When the information processing device 1000 receives the re-proposal request information, the extraction unit 1024 extracts an image to be newly provided to the user from the image information DB 1040 based on the user vector generated by the vector generation unit 1023. Here, text information "I want a stylish mug" is input in the character input area 1201 of the search screen 1200 (see FIG. 25). For this reason, as described in FIG. 25, it is possible to hold the category of an image already selected using a known selection technique in a memory (for example, the storage unit 1030) and execute a selection process to select an image to be provided to the electronic device 1100 from the image information DB 1040 based on the category of the image. That is, a selection process is executed to newly select an image to be provided to the electronic device 1100 from among categories (e.g., mugs) narrowed down in advance based on text information input in the character input area 1201 of the search screen 1200. This makes it possible to provide the user with appropriate image information according to the user's request. For example, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with tag information (feature amount) stored in the feature amount 1044 of the image information DB 1040, and based on the comparison result, newly extracts an image according to the user's preference. Specifically, the extraction unit 1024 calculates a difference value between each component constituting the user vector and each component constituting each tag information (each vector) for each tag information, and calculates a total value of the difference values for each tag information. Then, the extraction unit 1024 extracts an image corresponding to a predetermined number of tag information with a small total value of the difference values as a new image according to the user's preference. Note that, using the cosine similarity of each vector to be compared, an image corresponding to a predetermined number of tag information with a high cosine similarity may be extracted as a new image according to the user's preference. Also, other calculation methods may be used to extract new images according to the user's preferences.
[0249] Furthermore, the extraction unit 1024 outputs image information relating to the extracted new image (for example, image information 1042) to the information provision unit 1025.
[0250] [Example of selecting a new image] 28 is a diagram showing an example of a selection screen 1240 displayed on the UI section 1160 of the electronic device 1100. The selection screen 1240 is displayed on the UI section 1160 under the control of the control section 1150.
[0251] The selection screen 1240 displays an image display area 1241 that displays a plurality of images 1243 to 1250, a decision button 1251, and a re-suggest new information button 1252. Note that the selection screen 1240 is a display screen corresponding to the selection screen 1210 shown in Fig. 26, and the image display area 1241 and the decision button 1251 correspond to the image display area 1211 and the decision button 1221. Therefore, detailed explanations of these will be omitted. In addition, the selection operation and decision operation of each image are similar to the example shown in Fig. 26, and detailed explanations of these will be omitted.
[0252] Moreover, the re-propose new information button 1252 corresponds to the re-propose new information button 1233 shown in Fig. 27. When the re-propose new information button 1252 is selected, an image to be newly provided to the user (for example, an image other than the images displayed on the selection screen 1240) may be extracted from the image information DB 1040 based on the user vector and provided to the user, as in the example shown in Fig. 27, or an image extracted using other information may be provided to the user. Examples of extracting a new image using other information are shown in Figs. 36 to 40, etc.
[0253] Furthermore, after one or more images are selected by the user, when the user executes a selection operation (e.g., a touch operation) to select the decision button 1251, the control unit 1150 transmits image selection information related to the one or more images selected by the user to the information processing device 1000. Note that the tag generation process by the tag generation unit 1022, the vector generation process by the vector generation unit 1023, and the provider information extraction process by the extraction unit 1024 using the image selection information are similar to the above-mentioned examples.
[0254] In this way, it is possible to select a new image and provide it to the user using a user vector generated based on an image selected by the user (favorite image). In addition, by performing each of these processes multiple times, it is possible to refine recommendations to the user. In addition, a new image may be selected using a past calculation result (e.g., a user vector from one or several times ago). For example, it is possible to use a value obtained by sequentially adding up past calculation results and dividing the result by the number of calculation results to be added.
[0255] [Server operation example] Fig. 29 is a flowchart showing an example of the selection process by the information processing device 1000. This selection process is executed by the control unit 1020 (see Fig. 21) based on a program stored in the storage unit 1030 (see Fig. 21). This selection process is executed, for example, at the timing when an image request is received from the electronic device 1100. This selection process will be described with appropriate reference to Figs. 1 to 28.
[0256] In step S1301, the information providing unit 1025 transmits image information to the electronic device 1100 based on an image request from the electronic device 1100. This image information is information for displaying a selection screen 1210 (see FIG. 26) on the UI unit 1160 of the electronic device 1100. For example, the electronic device 1100 may be able to receive the service of the present embodiment by installing a predetermined application (search application) in the electronic device 1100. In this case, the control unit 1150 of the electronic device 1100 transmits an image request to the information processing device 1000 by starting the search application in the electronic device 1100. Also, as shown in FIG. 25, the selection screen 1210 (see FIG. 26) may be displayed in response to an input operation by a user. In this case, the electronic device 1100 transmits an image request including the input information (e.g., text information) to the information processing device 1000. Then, the extraction unit 1024 selects an image based on information (e.g., text information) included in the image request. For this selection process, a known image search method can be used.
[0257] In step S1302, the information acquisition unit 1021 determines whether or not image selection information (preferred image information) has been received from the electronic device 1100. For example, when the decision button 1221 is selected after the user has performed an image selection process on the selection screen 1210 (see FIG. 26), the control unit 1150 of the electronic device 1100 transmits the image selection information to the information processing device 1000. If the image selection information has been received, the process proceeds to step S1303. On the other hand, if the image selection information has not been received, monitoring is continued.
[0258] In step S1303, the tag generation unit 1022 acquires tag information associated with one or more images corresponding to the image selection information received in step S1302 from the image information DB 1040. Note that the method of acquiring this tag information is the same as the acquisition method described above.
[0259] In step S1304, the vector generation unit 1023 generates a user vector related to the user of the electronic device 1100 that transmitted the image selection information, based on the tag information acquired in step S1303. Note that the method of generating this user vector is the same as the above-mentioned generation method.
[0260] In step S1305, based on the user vector generated in step S1304, the extraction unit 1024 extracts and selects a provider that meets the preferences of the user corresponding to the user vector from among the providers stored in the provider information DB 1050. The provider selection method (extraction method) is the same as the extraction method described above.
[0261] In step S1306, the information providing unit 1025 transmits provider information (e.g., profile information 1052, image information 1042) related to the provider selected in step S1305 to the electronic device 1100. When the electronic device 1100 receives this provider information, the control unit 1150 of the electronic device 1100 causes the UI unit 1160 to display a provider information screen 1230 (see FIG. 27).
[0262] In step S1307, the information acquiring unit 1021 determines whether or not decision information has been received from the electronic device 1100 that transmitted the provider information in step S1306. For example, when the user selects the decision button 1234 on the provider information screen 1230 (see FIG. 27), the control unit 1150 of the electronic device 1100 transmits the decision information to the information processing device 1000. If decision information has been received, the process proceeds to step S1308. On the other hand, if decision information has not been received, the process proceeds to step S1309.
[0263] In step S1308, the control unit 1020 executes a predetermined process for introducing the provider corresponding to the determination information received in step S1307 (selected provider) to the user of the electronic device 1100. This predetermined process is similar to the predetermined process described above.
[0264] In step S1309, the information acquisition unit 1021 determines whether or not re-proposal request information has been received from the electronic device 1100 that transmitted the provider information in step S1306. For example, when a user selects the re-propose button 1233 for new information on the provider information screen 1230 (see FIG. 27), the control unit 1150 of the electronic device 1100 transmits the re-proposal request information to the information processing device 1000. When the re-proposal request information has been received, the process proceeds to step S1310. On the other hand, when the re-proposal request information has not been received, the process returns to step S1307.
[0265] In step S1310, based on the user vector generated in step S1304, the extraction unit 1024 extracts an image to be newly provided to the user from the image information DB 1040. This extraction process is similar to the extraction process described above (for example, the extraction process shown in FIG. 27).
[0266] In step S1311, the information providing unit 1025 transmits image information (for example, image information 1042) related to the new image extracted in step S1311 to the electronic device 1100. This image information is information for displaying the selection screen 1240 (see FIG. 28) on the UI unit 1160 of the electronic device 1100.
[0267] [Example of effect] Here, it is assumed that the user has an abstract ideal image and preferences regarding the desired information, but has difficulty in clearly verbalizing the image and desire. For this reason, if the user answers questions such as the image and desire of the desired information and information is proposed based on the user's answers, there is a risk that appropriate information cannot be proposed. Therefore, in this embodiment, it is possible to appropriately acquire the user's preferences by selecting an image according to the user's preferences from among a plurality of images displayed on the electronic device 1100. Then, a user vector is generated using tag information related to an image according to the user's preferences, and information (provider, image) is proposed to the user based on a comparison result between this user vector and a provider vector (or a vector of an image), so that it is possible to propose appropriate information (provider, image) according to the user's preferences. In other words, it is possible to support the user in acquiring the information desired intuitively and easily.
[0268] [Example of image editing by users] Here, it is assumed that there is no image provided to the user that matches the user's preferences. In this case, it is assumed that it is difficult for the user to select a favorite image. Therefore, it is conceivable that the image provided to the user can be temporarily edited to match the user's preferences, making it easier for the user to select a favorite image.
[0269] [Item information DB configuration example] 30 is a diagram showing, in a simplified form, the stored contents of the item information stored in the item information DB 1070. The item information DB 1070 may be stored in the storage unit 1030 of the information processing device 1000, or may be stored in an external device other than the information processing device 1000 and used.
[0270] The item information DB 1070 is a database for storing items for editing images provided by providers and feature amounts of a plurality of tags generated for the items in association with each other.
[0271] Specifically, image information 1072 , tag 1073 , and feature amount 1074 are stored in item information DB 1070 in association with item identification information 1071 .
[0272] The item identification information 1071 is identification information for identifying each item prepared by the service provider. Here, the service provider means a business operator that operates the information processing system IS1 and provides various services to users. Note that, for ease of explanation, FIG. 30 shows an example in which only the name indicating the item is stored in the item identification information 1071 as identification information, but other identification information may be used.
[0273] Image information 1072 is information about an item image prepared by a service provider so that the user can edit it. In Fig. 30, for ease of explanation, an example is shown in which only an image is stored in image information 1072, but various attribute information (e.g., size information) stored in association with the image may also be included.
[0274] The tag 1073 is information indicating a tag generated for an item image stored in the image information 1072. The feature amount 1074 is a feature amount assigned to the tag stored in the tag 1073. These tag information items are used when generating a user vector. Note that tag information having the same content as the feature amount 1044 shown in FIG. 23 may be stored, or tag information having different content from the feature amount 1044 shown in FIG. 23 may be stored. Also, only the necessary information among these items may be stored, and other information may be omitted.
[0275] [Example of image editing screen by user] 31 is a diagram showing an example of an edit screen 1260 displayed on the UI section 1160 of the electronic device 1100. The edit screen 1260 is displayed on the UI section 1160 under the control of the control section 1150.
[0276] The editing screen 1260 displays an image display area 1261 that displays a plurality of images 1262 to 1265, a decision button 1266, and an item display area 1270. Specifically, the information providing unit 1025 of the information processing device 1000 transmits editing screen information for displaying the editing screen 1260 to the electronic device 1100. When the electronic device 1100 receives the editing screen information, the control unit 1150 causes the UI unit 1160 to display the editing screen 1260 based on the received editing screen information. Note that the image display area 1261 corresponds to the image display area 1211 shown in FIG. 26.
[0277] The item display area 1270 is an area for displaying each of the item images 1271 to 1276 whose item information is stored in the item information DB 1070 (see FIG. 30). The user can generate an image of the user's preference by performing a moving operation to move a desired item displayed in the item display area 1270 to a desired image displayed in the image display area 1261. For example, when the UI unit 1160 is configured with a touch panel, the user can perform an operation (e.g., a drag-and-drop operation) to move a desired item while touching it to the position of the desired image. For example, the item image 1272 indicating a cactus can be arranged in the image 1262 by the user performing an operation to move the item image 1272 to a desired position in the image 1262. The transition of this movement is shown typically by an arrow 1277. The image in which the item image is arranged (edited image) is set to a selected state. The image synthesis process in which the item image is arranged in the image and synthesized can use a known image processing. In addition, various image processes (e.g., zooming in and out, adjusting the three-dimensional angle, additional editing of the image) may be performed on each item image based on a user operation. For each of these image processes, known image processing may be used. In addition, the item image arranged in the image may be appropriately deformed in relation to other components in the image using known image processing techniques. For example, it is possible to zoom in and out on the item image based on the size of each object in the image, or to adjust the angle of the item image based on the inclination of each object in the image.
[0278] In this way, when an image matching the user's preference is not displayed in the image display area 1261, it is possible to generate an image matching the user's preference by editing a part of each image displayed in the image display area 1261. Note that the image generated by this editing is an image temporarily generated for selecting the provider or a new image, and is deleted after the provider or new image selection process is completed.
[0279] In this way, when a user operation is accepted by the accepting unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. For example, when a selection operation (e.g., a touch operation) for selecting an image to be edited from among the images displayed in the image display area 1261 is performed by the user, the preferred image selected by the selection operation is placed in a selected state.
[0280] Furthermore, when the user executes a selection operation (e.g., a touch operation) of selecting the decision button 1266 after editing one or more images, the control unit 1150 transmits edited image information regarding the one or more images edited by the user to the information processing device 1000. When the information processing device 1000 receives the edited image information, the information acquisition unit 1021 outputs the received edited image information to the tag generation unit 1022 and the vector generation unit 1023.
[0281] [Example of generating user vectors] The tag generation unit 1022 acquires tag information for one or more edited images corresponding to the edited image information transmitted from the electronic device 1100 from the image information DB 1040 and the item information DB 1070. Specifically, the tag generation unit 1022 acquires tag information associated with the original image of each edited image from the image information DB 1040, and acquires tag information associated with an item added to each edited image by editing from the item information DB 1070. Then, the tag generation unit 1022 outputs each acquired tag information to the vector generation unit 1023.
[0282] Next, the vector generation unit 1023 generates a user vector related to the user who transmitted the edited image information, based on the tag information output from the tag generation unit 1022. In this case, the method of generating the user vector is the same as that described above, except that the tag information associated with the item is used, and therefore detailed description here will be omitted.
[0283] Here, it is considered that the user's preference is reflected more in the edited image in which an item is added by the user than in the image that is not edited. Therefore, the user vector may be generated by weighting the tag information associated with the edited image in which an item is added by the user more than other tag information. It is also assumed that the edited image has a plurality of items added. In the case of such an edited image, it is considered that the user's preference is further reflected. Therefore, for example, the user vector may be generated by increasing the weighting for the edited image and the weighting for the added items according to the number of items added to the edited image. That is, the user vector may be generated by changing the weighting for the edited image and the weighting for the added items based on various information related to the items added to the edited image (for example, the number of items, the ratio of item images in the image). This makes it possible to generate a user vector that reflects the user's preference.
[0284] Here, as shown in FIG. 31, when an image is edited by adding an item to the image, it is assumed that the design taste of the object (e.g., a mug) included in the image may differ depending on the position of the item in the object. Therefore, using the above-mentioned AI model, the tag generating unit 1022 may generate new tag information for the edited image. In this case, it is possible to generate more appropriate tag information according to the editing result edited by the user.
[0285] [Examples of electronic device operation] Fig. 32 is a flowchart showing an example of image editing processing by electronic device 1100. This image editing processing is executed by control unit 1150 (see Fig. 21) based on a program stored in storage unit 1170 (see Fig. 21). This image editing processing is executed, for example, at the timing when a user operation is performed to request an editing screen. This image editing processing will be described with appropriate reference to Figs. 1 to 31.
[0286] In step S1321, the control unit 1150 transmits editing screen request information for requesting editing screen information to the information processing device 1000. For example, when a user operation for requesting an editing screen is performed, the control unit 1150 transmits the editing screen request information to the information processing device 1000.
[0287] In step S1322, the control unit 1150 determines whether or not editing screen information has been received from the information processing device 1000. If editing screen information has been received, the process proceeds to step S1323. On the other hand, if editing screen information has not been received, monitoring is continued.
[0288] In step S1323, the control unit 1150 causes the UI unit 1160 to display an editing screen (for example, the editing screen 1260 (see FIG. 31)) based on the editing screen information received in step S1322.
[0289] In step S1324, the control unit 1150 determines whether or not a user operation on the editing screen displayed in step S1323 has been accepted by the accepting unit 1161. If a user operation has been accepted, the process proceeds to step S1325. On the other hand, if a user operation has not been accepted, the process continues to monitor.
[0290] In step S1325, the control unit 1150 determines whether or not the user operation accepted in step S1324 is an editing operation for editing the image displayed on the editing screen. For example, when a moving operation is performed to move a desired item displayed in the item display area 1270 (see FIG. 31) to a desired image displayed in the image display area 1261 (see FIG. 31), it is determined that an editing operation has been performed as the user operation. If the user operation is an editing operation, the process proceeds to step S1326. On the other hand, if the user operation is not an editing operation, the process proceeds to step S1327.
[0291] In step S1326, the control unit 1150 executes an editing process to edit the image based on the user operation (editing operation) accepted in step S1324. For example, when a move operation is performed to move an item displayed in the item display area 1270 to the image displayed in the image display area 1261, image processing is executed to move the item in response to the move operation and combine it with the image. Also, for example, when a delete operation is performed to delete the item moved to the image, image processing is executed to delete the item in response to the delete operation.
[0292] In step S1327, the control unit 1150 determines whether the user operation accepted in step S1324 is a decision operation that is performed after one or more images are edited. For example, when a selection operation (e.g., a touch operation) for selecting the decision button 1266 (see FIG. 31) is performed, it is determined that the decision operation is performed as the user operation. If the user operation is the decision operation, the process proceeds to step S1329. On the other hand, if the user operation is not the decision operation, the process proceeds to step S1328.
[0293] In step S1328, the control unit 1150 executes a predetermined process based on the user operation accepted in step S1324. For example, when a selection operation is performed to place the image displayed in the image display area 1261 in a selected state, image processing is executed to place the image in a selected state in response to the selection operation.
[0294] In step S1329, the control unit 1150 transmits to the information processing device 1000 edited image information relating to one or more images edited by the user.
[0295] In this way, by making it possible to edit an image according to the user's preferences from among a plurality of images displayed on electronic device 1100, it is possible to more appropriately express the user's preferences based on the user's subjective editing. Therefore, it is possible to propose appropriate information (providers, images) according to the user's preferences based on the user's editing results. In other words, it is possible to support the user in obtaining desired information in an intuitive and easy-to-understand manner.
[0296] In the above, an example of editing an image using a prepared image (item image) has been shown, but an image may be edited using text information input by a user (for example, manual input, voice input, gesture input). For example, it is possible to input "I like piano" or "I like sofa" as text information. In this case, for example, it is possible to reflect the input text information on the image in the selected state. For example, it is possible to generate an image related to the input text information using a generative AI (for example, an image generation AI (for example, Stable Diffusion, Midjourney)) that has learned about each tag. Then, it is possible to generate a desired edited image by synthesizing the generated image (image related to text information) with the image in the selected state. In this case, for example, a known image synthesis technique can be used. This makes it possible to generate a desired edited image related to the input text information. In this case, for example, a user vector may be generated by increasing the weight associated with the edited image and the weight associated with the input text information according to the weight of the input text information (for example, the number of characters). This makes it possible to generate a user vector that reflects the user's preferences.
[0297] [Example using user-acquired images] As described above, it is possible that the images provided to the user do not match the user's preferences. In this case, it is possible that the user has difficulty selecting a preferred image. Therefore, it is possible to select information according to the user's preferences by generating a user vector according to the user's preferences using images acquired by the user.
[0298] [Image information DB configuration example] FIG. 33 is a simplified diagram showing the contents of the image information stored in the image information DB 1080. As shown in FIG.
[0299] The image information DB 1080 is a database stored in the storage unit 1170 (see FIG. 21) of the electronic device 1100 or an external device (for example, a server). The image information DB 1080 also stores image information acquired by the image acquisition unit 1130. For example, when a smartphone is used as the electronic device 1100, a camera mounted on the smartphone corresponds to the image acquisition unit 1130.
[0300] When a user performs an image acquisition operation (for example, an operation to shoot a still image or an operation to shoot a video), the control unit 1150 of the electronic device 1100 records the image acquired by the image acquisition unit 1130 based on the acquisition operation as an image file (still image file, video file) in the image information DB 1080. For example, as shown in Fig. 33, the image file is stored in the image information 1082 in association with the image identification information 1081.
[0301] In this case, the control unit 1150 associates the location information (e.g., latitude and longitude) acquired by the location information acquisition unit 1120 of the electronic device 1100 with the image file and stores it based on the timing of the acquisition operation. For example, as shown in Fig. 33, the acquired location information is stored in the location information 1083 in association with the image file. For example, when a smartphone is used as the electronic device 1100, a GPS device mounted on the smartphone corresponds to the location information acquisition unit 1120.
[0302] Furthermore, the control unit 1150 stores the audio information acquired by the audio acquisition unit 1140 of the electronic device 1100 in association with the image file based on the timing of the acquisition operation. For example, as shown in FIG. 33, the acquired audio information is stored in the audio information 1084 in association with the image file. For example, when a smartphone is used as the electronic device 1100, a microphone mounted on the smartphone corresponds to the audio acquisition unit 1140. When a still image is shot, audio information of a predetermined time based on the shooting timing can be stored in association with the still image. Furthermore, when a video is shot, the acquired audio information may be used to use a representative image of the video (for example, the first frame, the frame with the highest feature) as a still image.
[0303] For example, a user photographs various subjects, scenery, and the like that suit the user's taste using the electronic device 1100, and the images that suit the user's taste are stored in the image information DB 1080. Therefore, it is possible to generate a user vector according to the user's taste using the image information and the like stored in the image information DB 1080. When a stationary device such as a personal computer is used as the electronic device 1100, it is possible that the position information acquisition unit 1120 and the image acquisition unit 1130 are not provided. In this case, an external device may be used as the position information acquisition unit 1120 and the image acquisition unit 1130, and image information and the like acquired using an imaging device (for example, a smartphone, a digital still camera, a digital video camera (for example, a camera-integrated recorder)) is acquired and stored in the image information DB 1080, and this can be used.
[0304] [Example of user selection of preferred image] 34 is a diagram showing an example of a selection screen 1280 displayed on the UI section 1160 of the electronic device 1100. The selection screen 1280 is displayed on the UI section 1160 under the control of the control section 1150.
[0305] The selection screen 1280 displays an image display area 1290 for displaying a plurality of images 1291 to 1296, a selected image display area 1281 for displaying a selected image, a position information designation area 1283, an audio information designation area 1284, and a decision button 1285. Specifically, the information providing unit 1025 of the information processing device 1000 transmits selection screen information for displaying the selection screen 1280 to the electronic device 1100. When the electronic device 1100 receives the selection screen information, the control unit 1150 causes the UI unit 1160 to display the selection screen 1280 based on the received selection screen information.
[0306] The image display area 1290 is an area for displaying images 1291 to 1296 whose image information is stored in the image information DB 1080 of the electronic device 1100. That is, the control unit 1150 acquires each piece of image information stored in the image information DB 1080, and displays it in the image display area 1290 of the selection screen 1280. Note that the images to be displayed here are not particularly limited, and various types of images can be displayed.
[0307] The user can select a desired image by performing a moving operation to move a desired image displayed in the image display area 1290 to the selected image display area 1281. For example, when the UI unit 1160 is configured with a touch panel, the user can perform an operation (e.g., a drag-and-drop operation) to move the desired image to the selected image display area 1281 while touching it. For example, the user can place the image 1292 in the selected image display area 1281 by performing an operation to move the image 1292 including a mug to the selected image display area 1281. The transition of this movement is shown typically by an arrow 1297. Note that the image already displayed in the selected image display area 1281 and another image may be composited using a known image processing technique to create a composite image. Also, one or more images may be edited using each of the above-mentioned editing techniques.
[0308] The selected image display area 1281 is an area for displaying the images 1218, 1292 selected based on a user operation. The selected image display area 1281 may display the images selected on the screens shown in Fig. 26 and Fig. 28, the edited images edited on the edit screen 1260 shown in Fig. 31, and the like. When various images are displayed in this manner, it may be possible to appropriately perform operations for deleting, adding, and editing the various images based on a user operation.
[0309] FIG. 34 shows an example in which the image 1218 selected in FIG.
[0310] The position information designation area 1283 is an area operated when designating whether or not to use the position information associated with the image displayed in the image display area 1290. For example, when a selection operation of selecting the pull-down button of the position information designation area 1283 is performed after the image moved to the selected image display area 1281 is selected, a selection area for selecting use or non-use is displayed, and the user performs the operation desired in this selection area. For example, when designating to use the position information associated with the selected image, it is possible to generate a user vector using the position information associated with the image. On the other hand, when designating not to use the position information associated with the selected image, it is possible to generate a user vector without using the position information associated with the image.
[0311] The audio information designation area 1284 is an area operated when designating whether or not to use audio information associated with an image displayed in the image display area 1290. For example, when a selection operation of selecting a pull-down button in the audio information designation area 1284 is performed after an image moved to the selected image display area 1281 is selected, a selection area for selecting use or non-use is displayed, and the user performs a desired operation in this selection area. For example, when designating to use audio information associated with the selected image, it is possible to generate a user vector using the audio information associated with the image. On the other hand, when designating not to use audio information associated with the selected image, it is possible to generate a user vector without using the audio information associated with the image.
[0312] In this way, when a user operation is accepted by the accepting unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. In addition, when the user executes a selection operation (e.g., a touch operation) of selecting the decision button 1285 after moving one or more images to the selected image display area 1281, the control unit 1150 transmits selected image information related to one or more images selected by the user to the information processing device 1000. This selected image information includes image information, position information (if use is specified), and audio information (if use is specified). When the information processing device 1000 receives the selected image information, the information acquisition unit 1021 outputs the received selected image information to the tag generation unit 1022 and the vector generation unit 1023.
[0313] [Example of generating user vectors] The tag generating unit 1022 generates tag information for one or more selected images corresponding to the selected image information transmitted from the electronic device 1100. When an image stored in the image information DB 1040 (see FIG. 23) is selected as the selected image, the tag generating unit 1022 obtains the tag information of the image from the image information DB 1040.
[0314] Furthermore, if an image stored in the image information DB 1080 of the electronic device 1100 is a selected image, the tag generation unit 1022 generates tag information for that image. For example, it is possible to generate tag information for that image using the above-mentioned AI model.
[0315] In addition, when at least one of the location information (when use is specified) and the voice information (when use is specified) is included in the selected image information, it is possible to generate tag information for each piece of information. For example, a database that associates location information (e.g., latitude and longitude) with tag information (e.g., information previously set for each region) is prepared in advance in the storage unit 1030 of the information processing device 1000, and tag information corresponding to the location information included in the selected image information can be obtained using this database.
[0316] Also, for example, a database that associates location information (e.g., latitude and longitude) with the characteristics and environmental images of a certain location (e.g., Okinawa is "hot" and Kamakura is "ancient capital" and "calm") can be prepared in advance in the storage unit 1030 of the information processing device 1000, and this database can be used to obtain the characteristics and environmental images (e.g., Okinawa is "hot" and Kamakura is "ancient capital" and "calm") corresponding to the location information included in the selected image information, and text information related to these characteristics and environmental images (e.g., "hot", "ancient capital", "calm"). In this case, tag information can be generated from the text information using an AI model described later.
[0317] In addition, for example, a database that associates audio information with tag information (for example, information pre-set for each type of sound (for example, "noise," "quiet," "river babbling," "children's voices")) can be prepared in advance in the memory unit 1030 of the information processing device 1000, and this database can be used to obtain tag information corresponding to the audio information included in the selected image information.
[0318] Also, for example, it is possible to convert voice information into text information. In this case, it is possible to generate tag information from the text information using an AI model described later. Also, tag information may be generated using an AI model that has been trained using location information, voice information, and the like.
[0319] Next, the vector generation unit 1023 generates a user vector related to the user who transmitted the image selection information, based on the tag information output from the tag generation unit 1022. Note that the method of generating this user vector is similar to the method of generating the user vector described above, and therefore a detailed description thereof will be omitted here.
[0320] In the above, an example has been shown in which tag information is generated using an AI model that has been previously learned and generated, and a user vector is generated using this tag information, but other AI models may be used. For example, a large language model (LLM) may be used as the AI model. For example, ChatGPT (Generative Pre-trained Transformer), Bard, Llama2 (Large Language Model Meta AI 2), Gemini, Claude, etc. may be used as the LLM. Note that these are only examples, and other AI models may be used. In addition, it is also possible that text information is attached to an image (e.g., an advertisement image). In such a case, tag information related to the text information may be generated using an AI model, and the tag information obtained by combining this tag information with tag information related to an image (e.g., an advertisement image) may be used to generate a user vector.
[0321] [Examples of electronic device operation] Fig. 35 is a flowchart showing an example of image selection processing by electronic device 1100. This image selection processing is executed by control unit 1150 (see Fig. 21) based on a program stored in storage unit 1170 (see Fig. 21). This image selection processing is executed, for example, at the timing when a user operation is performed to request a selection screen. This image selection processing will be described with appropriate reference to Figs. 1 to 34.
[0322] In step S1331, the control unit 1150 transmits selection screen request information for requesting selection screen information to the information processing device 1000. For example, when a user operation for requesting a selection screen is performed, the control unit 1150 transmits the selection screen request information to the information processing device 1000.
[0323] In step S1332, the control unit 1150 determines whether or not selection screen information has been received from the information processing device 1000. If selection screen information has been received, the process proceeds to step S1333. On the other hand, if selection screen information has not been received, monitoring is continued.
[0324] In step S1333, control unit 1150 acquires, from image information DB 1080 (see FIG. 33), image information (for example, image information 1082) to be displayed on the editing screen corresponding to the selection screen information received in step S1332.
[0325] In step S1334, the control unit 1150 displays a selection screen (for example, the selection screen 1280 (see FIG. 34)) on the UI unit 1160 based on the selection screen information received in step S1332. The image information acquired in step S1333 is displayed on this selection screen.
[0326] In step S1335, the control unit 1150 determines whether or not a user operation on the selection screen displayed in step S1334 has been accepted by the acceptance unit 1161. If a user operation has been accepted, the process proceeds to step S1336. On the other hand, if a user operation has not been accepted, the process continues to monitor.
[0327] In step S1336, the control unit 1150 determines whether the user operation accepted in step S1335 is a selection operation for selecting an image displayed on the selection screen. For example, when a move operation is performed to move an image displayed in the image display area 1290 (see FIG. 34) to the selected image display area 1281 (see FIG. 34), it is determined that a selection operation has been performed as the user operation. If the user operation is a selection operation, the process proceeds to step S1337. On the other hand, if the user operation is not a selection operation, the process proceeds to step S1338.
[0328] In step S1337, the control unit 1150 executes a selection process for selecting an image based on the user operation (selection operation) accepted in step S1335. For example, when a move operation is performed to move an image displayed in the image display area 1290 (see FIG. 34) to the selected image display area 1281 (see FIG. 34), image processing is executed for moving the image in response to the move operation to set it as a selected image. Also, for example, when a delete operation is performed to delete an image moved to the selected image display area 1281 as the selected image, image processing is executed for deleting the image in response to the delete operation from the selected image display area 1281.
[0329] In step S1338, the control unit 1150 determines whether the user operation accepted in step S1335 is a decision operation that is performed after one or more images are selected. For example, when a selection operation (e.g., a touch operation) for selecting the decision button 1285 (see FIG. 34) is performed, it is determined that the decision operation is performed as the user operation. If the user operation is the decision operation, the process proceeds to step S1340. On the other hand, if the user operation is not the decision operation, the process proceeds to step S1339.
[0330] In step S1339, the control unit 1150 executes a predetermined process based on the user operation accepted in step S1335. For example, when a selection operation is performed to select the pull-down button of the position information designation area 1283 or the audio information designation area 1284 (see FIG. 34), image processing is executed to use or not use the corresponding area in accordance with the selection operation.
[0331] In step S1340, the control unit 1150 transmits to the information processing device 1000 selected image information relating to one or more images selected by the user.
[0332] In this way, by allowing the user to select an image according to the user's preference from among images acquired by using electronic device 1100, it is possible to more appropriately express the user's preference based on the user's subjective image collection. Therefore, it is possible to propose appropriate information (providers, images) according to the user's preference based on the user's image collection behavior (e.g., shooting behavior). In other words, it is possible to support the user in acquiring desired information in an intuitive and easy-to-understand manner.
[0333] [Example using text information] In the above, an example of selecting a provider, a new image, etc., mainly using an image has been shown. However, it is assumed that a provider, a new image, etc. that the user prefers may not be presented to the user. Therefore, in the following, an example of selecting a provider, a new image, etc., using text information will be shown.
[0334] [Example of generating user vectors based on text information] 36 is a diagram showing an example of a text input screen 1400 displayed on the UI section 1160 of the electronic device 1100. The text input screen 1400 is displayed on the UI section 1160 under the control of the control section 1150.
[0335] The text input screen 1400 displays an image display area 1401 that displays a plurality of images 1402 to 1404, a text information input field 1405, a re-suggest button 1406, and a decision button 907. Specifically, the information providing unit 1025 of the information processing device 1000 transmits screen information for displaying the text input screen 1400 to the electronic device 1100 based on a request from the electronic device 1100. When the electronic device 1100 receives the screen information, the control unit 1150 causes the UI unit 1160 to display the text input screen 1400 based on the received screen information. It is possible to display other images on the text input screen 1400 based on a user operation.
[0336] Furthermore, when a user operation is received by the receiving unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. Here, it is assumed that the image display area 1401 does not display the user's favorite image. In such a case, it is possible for the user to input (for example, manual operation, voice input) characters related to the favorite image (i.e., the product or service of the image) into the text information input field 1405. FIG. 36 shows an example in which the user inputs "I prefer a slightly more modern design." Note that this is just one example, and characters related to the favorite image may be, for example, "I like something a little simpler," "I like a cuter design," "I want something with an ethnic nuance," etc.
[0337] Furthermore, when the user executes a selection operation (for example, a touch operation) to select the re-suggest button 1406 after characters are input in the text information input field 1405, the control unit 1150 transmits text information corresponding to the characters input by the user to the information processing device 1000. When the information processing device 1000 receives the text information, the information acquisition unit 1021 outputs the received text information to the tag generation unit 1022 and the vector generation unit 1023. Note that, when one or more images are selected by the user, the control unit 1150 transmits image selection information related to the selected one or more images together with the text information to the information processing device 1000. Also, in this case, a category (for example, mug) narrowed down in advance based on the text information (for example, I want a stylish mug) input in the character input area 1201 of the initial search screen (for example, search screen 1200 (see FIG. 25)) is held in the memory of the information processing device 1000. Therefore, a selection process is executed to newly select an image to be provided to the electronic device 1100 based on the category held in the memory.
[0338] [Example of generating user vectors] The tag generating unit 1022 generates tag information corresponding to the text information transmitted from the electronic device 1100. When one or more images are selected by the user, image selection information regarding the selected image or images is transmitted from the electronic device 1100 together with the text information. In this case, it is possible to generate a user vector taking into consideration the image selection information. Examples of this are shown in Figs. 39 and 40.
[0339] For example, the tag generating unit 1022 can generate tag information corresponding to text information by using a predetermined database (for example, a dictionary database) for converting text information into tag information. For example, a dictionary database in which character strings and features of each element constituting tag information (for example, each feature corresponding to the tag 1043 shown in FIG. 23) are associated with each other can be used. Then, the tag generating unit 1022 extracts character strings included in the dictionary database from each character string included in the sentence corresponding to the text information transmitted from the electronic device 1100, and extracts features of tag information corresponding to the extracted character strings from the dictionary database. A known character recognition technique can be adopted as a method for extracting this character string. Also, when multiple character strings are extracted, features of tag information corresponding to each character string are extracted. Then, the tag generating unit 1022 counts up (for example, calculates an average value) the features of one or more pieces of extracted tag information to generate tag information corresponding to the text information transmitted from the electronic device 1100. Then, the tag generating unit 1022 outputs the generated tag information to the vector generating unit 1023.
[0340] Also, it is possible to generate tag information (user vector) from text information using an AI model. For example, it is possible to use an AI model that has learned a plurality of pieces of text information to which predetermined tags that associate a sentence with the features of each element constituting the tag information (for example, each feature corresponding to the tag 1043 shown in FIG. 23) are added as teacher labels. As for this learning method, it is possible to apply the learning method shown in FIG. 22.
[0341] Also, for example, LLM can be used as the AI model. For example, ChatGPT, Bard, Llama2, Gemini, Claude, etc. can be used as the LLM. Note that these are only examples, and other AI models may be used. For example, based on text information input by a user, condition information to output features related to an image for each of a plurality of tags can be input to the LLM as instruction information (prompt), and the output information in response to this can be tag information.
[0342] Here, an example of a prompt that suggests tags using ChatGPT is shown. For example, by using a sentence constituting the text information transmitted from the electronic device 1100, instruction information (prompt) to output a score (0 to 1) regarding the relevance with each of a plurality of tags (for example, tag 1043 shown in FIG. 23) can be input to the LLM, and the output information in response to this can be set as the feature amount for each of the plurality of tags. Note that the tags are not limited to those shown in FIG. 23, and other tags may be used. For example, Western style, Japanese style, modern, American, antique, country, simple, natural, Nordic, industrial, woody, family gathering, warmth, high-class, elegant, luxury, cool, chic, casual, pop, refined, stylish, open feeling, calm, flashy, profound feeling, cute, etc. may be used as tag elements.
[0343] Next, the vector generation unit 1023 generates a user vector related to the user who transmitted the text information, based on the tag information output from the tag generation unit 1022. In this case, the tag information generated by the tag generation unit 1022 can be used as the user vector. Note that, when there are multiple pieces of text information transmitted from the electronic device 1100, tag information is generated for each piece of text information, so that it is possible to generate a user vector using multiple pieces of tag information, similar to the example shown in FIG.
[0344] [Example of image selection based on text information] Next, the extraction unit 1024 extracts an image to be newly provided to the user from the image information DB 1040 based on the user vector generated by the vector generation unit 1023. This extraction method is similar to the method of extracting a new image when the "re-suggest new information" button 1233 (see FIG. 27) is selected, and therefore will not be described here. For example, a new image extraction process is performed using a category (e.g., mugs) stored in the memory of the information processing device 1000.
[0345] Furthermore, the extraction unit 1024 outputs image information (for example, image information 1042) relating to the extracted image to the information providing unit 1025. Then, the information providing unit 1025 transmits the image information to the electronic device 1100, and the electronic device 1100 displays the image information.
[0346] [Example of selecting a new image] 37 is a diagram showing a display example of a selection screen 1410 displayed on the UI unit 1160 of the electronic device 1100. The selection screen 1410 is displayed on the UI unit 1160 based on the control of the control unit 1150. The selection screen 1410 is a modified version of a part of the message at the top of the selection screen 1240 shown in FIG. 28, and the rest is common to the selection screen 1240. For this reason, the same reference numerals are used to denote the parts common to the selection screen 1240. The image display area 1241 displays each image selected by the image selection process based on the text information described above, but here, for ease of explanation, an example is shown in which a plurality of images 1243 to 1250 similar to the selection screen 1240 are displayed.
[0347] In this way, it is possible to select a new image and provide it to the user using the user vector generated based on the text information input by the user. Note that a text information input field 1405 (see FIG. 36) may be provided in the selection screen 1410 to allow further input of text information. In this case, when the text information is input by the user, it is possible to re-suggest a new image to the user based on the text information. Furthermore, by performing each of these processes multiple times, it is possible to refine the recommendation to the user.
[0348] In the above, an example has been shown in which text information is accepted by a user input when there is no preferred image among the images presented to the user, when there is a request for the images, etc. However, this is not limited to this. For example, as in the example shown in FIG. 25, it is possible to accept text information by a user input before an image is presented, and to select an image and present it to the user using a user vector generated based on this text information.
[0349] [Server operation example] Fig. 38 is a flowchart showing an example of the selection process by the information processing device 1000. This selection process is a partial modification of the selection process shown in Fig. 29. Specifically, steps S1351 to S1354 are added. Other than this addition, the selection process is the same as that shown in Fig. 29, so the same reference numerals are used to designate the same parts as in Fig. 29, and the description thereof will be omitted.
[0350] In step S1351, the information acquiring unit 1021 determines whether or not text information has been received from the electronic device 1100. For example, when a user performs an input process of text information on the text input screen 1400 (see FIG. 36) and then selects the re-suggest button 1406, the control unit 1150 of the electronic device 1100 transmits the text information to the information processing device 1000. If the text information has been received, the process proceeds to step S1352. On the other hand, if the text information has not been received, the process returns to step S1302.
[0351] In step S1352, the tag generation unit 1022 and the vector generation unit 1023 generate tag information (user vector) based on the text information received in step S1351. The method of generating this user vector is similar to the above-mentioned generation method.
[0352] In step S1353, the extraction unit 1024 extracts, from the image information DB 1040, an image to be newly provided to the user, based on the user vector generated in step S1352.
[0353] In step S1354, the information providing unit 1025 transmits image information (for example, the image information 1042) relating to the new image extracted in step S1353 to the electronic device 1100. This image information is information for displaying the selection screen 1410 (see FIG. 37) on the UI unit 1160 of the electronic device 1100.
[0354] [Example of image re-recommendation based on image and text information] In the above, an example of presenting an image to a user using a user vector generated based on text information has been shown. Here, after presenting a plurality of images to a user, it is assumed that an image that the user likes is selected from among them, and a request from the user regarding the images is received by the user through text information. Therefore, in the following, an example of generating a new user vector based on an image selected by the user and a request from the user regarding the image presented to the user will be shown.
[0355] [Example of generating user vectors based on selected images and text information] Fig. 39 is a diagram showing a display example of a text input screen 1430 displayed on the UI unit 1160 of the electronic device 1100. Note that the text input screen 1430 is similar to the text input screen 1400 shown in Fig. 36, and differs in that the content input in the text information input field 1405 and the images 1402 and 1403 are in a selected state (star regions are added), and the rest is common to the text input screen 1400. For this reason, the parts common to the text input screen 1400 are shown with the same reference numerals. In addition, the transmission process of each piece of information to the information processing device 1000 and the like are also similar to the example shown in Fig. 36.
[0356] As described above, it is possible that the image display area 1401 does not display the user's preferred image, that there are few images that the user likes, or that an image slightly different from the user's preferred image is displayed. In such cases, it is possible for the user to input text related to the preferred image in the text information input field 1405 by inputting (for example, manual operation or voice input). FIG. 39 shows an example in which the user has input "Please recommend a mug with a slightly brighter atmosphere." Note that this is just one example, and text related to the preferred image may be, for example, "Please make the original recommendation result blank and recommend an image of a mug with a brighter atmosphere."
[0357] [Example of generating user vectors] The tag generating unit 1022 generates tag information corresponding to each of the image selection information and the text information transmitted from the electronic device 1100. The generation of tag information based on the image selection information is similar to the example shown in Fig. 26 and the like. The generation of tag information based on the text information is similar to the example shown in Fig. 36. The generation of user vectors based on the tag information generated by the tag generating unit 1022 (user vectors based on the image selection information, user vectors based on the text information) is also similar to each of the generation examples described above.
[0358] [Example of determining the ratio of image selection information and text information based on user requests] Here, an example is shown in which a user vector C for selecting a new image is generated using a user vector (user vector A) generated based on image selection information and a user vector (user vector B) generated based on text information.
[0359] For example, it is possible to generate user vector C using a preset ratio (reflection ratio of user vector A and user vector B). For example, it is possible to obtain user vector C by the following formula 7 using reflection ratio α of user vector A and user vector B. Note that the calculation method according to formula 7 is only one example, and user vector C may be obtained by other calculation methods. User vector C = (1-α) × User vector A + α × User vector B … Equation 7
[0360] Furthermore, the vector generation unit 1023 generates a user vector A and a user vector B based on the tag information output from the tag generation unit 1022, and generates a user vector C based on the reflection ratio α and the user vector A and the user vector B. The reflection ratio α may be set by the user or may be set by the service provider in this embodiment.
[0361] The reflection ratio α may be set based on at least one of the text information. For example, it is possible to set the reflection ratio α using an AI model. For example, it is possible to use an AI model that has learned a plurality of text information to which a predetermined tag associated with the reflection ratio and the request sentence is added as a teacher label.
[0362] Also, for example, LLM can be used as the AI model. For example, ChatGPT, Bard, Llama2, Gemini, Claude, etc. can be used as the LLM. Note that these are only examples, and other AI models may be used. For example, based on text information input by a user, condition information for outputting the reflection ratio α can be input to the LLM as instruction information (prompt), and the output information in response to this can be the reflection ratio α.
[0363] Here, an example of a prompt for outputting the reflection ratio α using ChatGPT is shown. For example, it is possible to input instruction information (prompt) to the LLM to determine the extent to which the content of a sentence constituting the text information transmitted from the electronic device 1100 should be reflected, determine the ratio (0 to 1), and output the instruction information (prompt), and set the output information to the reflection ratio α. In this case, instruction information (prompt) to output the estimated basis for the reflection ratio α may be input to the LLM.
[0364] For example, as an example of how to determine the reflection ratio α (example of the estimation basis of the reflection ratio α), when a user inputs "Please blank out the original recommendation result and recommend an image of a mug with a bright atmosphere," it is possible to estimate the reflection ratio α by considering that the reflection ratio of the content of the text is appropriate to be 1.0, since the request to ignore the original recommendation result is seen. Also, for example, when a user inputs "Please recommend a mug with a brighter atmosphere," it is possible to estimate the reflection ratio α by considering that the reflection ratio of the content of the text is appropriate to be about 0.3, since the request to maintain the original recommendation result to some extent and change it a little more is seen. Also, for example, when a user inputs "I like mugs with a bright atmosphere," it is difficult to determine whether the original recommendation result should be maintained, so it is possible to estimate the reflection ratio α by considering that the reflection ratio of the content of the text is about 0.5.
[0365] [Image selection example] Next, the extraction unit 1024 extracts an image to be newly provided to the user from the image information DB 1040 based on the user vector C generated by the vector generation unit 1023. This extraction method is similar to the method of extracting a new image when the "re-suggest new information" button 1233 (see FIG. 27) is selected, and therefore will not be described here. For example, a new image extraction process is performed using a category (e.g., mugs) stored in the memory of the information processing device 1000.
[0366] Furthermore, the extraction unit 1024 outputs image information (for example, image information 1042) relating to the extracted image to the information providing unit 1025. Then, the information providing unit 1025 transmits the image information to the electronic device 1100, and the electronic device 1100 displays the image information.
[0367] In the above, an example has been shown in which an image to be newly provided to the user is extracted based on the user vector C, but a provider according to the user's preference may be extracted from the provider information DB 1050 based on the user vector C. In this case, provider information (e.g., profile information 1052, image information 1042) related to the extracted provider is transmitted to the electronic device 1100 and displayed on the electronic device 1100. For example, a provider information screen 1230 (see FIG. 27) is displayed on the UI unit 1160.
[0368] Also, a text information input field 1405 (see FIGS. 36 and 39) may be provided on the provider information screen 1230 (see FIG. 27), and a user vector B may be generated according to the user's preferences using text information inputted into the text information input field 1405, and an image extraction process, a provider extraction process, and the like may be executed using the already generated user vector A (user vector generated based on image selection information) and user vector B. In this case, each piece of extracted information is transmitted to the electronic device 1100 and displayed on the electronic device 1100.
[0369] [Example of selecting a new image] 40 is a diagram showing a display example of a selection screen 1440 displayed on the UI unit 1160 of the electronic device 1100. The selection screen 1440 is displayed on the UI unit 1160 based on the control of the control unit 1150. The selection screen 1440 is a modified version of a part of the message at the top of the selection screen 1240 shown in FIG. 28, and the rest is common to the selection screen 1240. For this reason, the same reference numerals are used to denote the parts common to the selection screen 1240. The image display area 1241 displays each image selected by the image selection process based on the user vector C described above, but here, for ease of explanation, an example is shown in which a plurality of images 1243 to 1250 similar to the selection screen 1240 are displayed.
[0370] In this way, it is possible to select a new image and provide it to the user using the user vector C generated based on the text information input by the user and the image selected by the user. Note that a text information input field 1405 (see FIG. 36 and FIG. 39) may be provided in the selection screen 1440 to allow further input of text information. In this case, when the text information is input by the user, it is possible to provide the user with a further new image based on the text information. Furthermore, by performing each of these processes multiple times, it is possible to refine the recommendations to the user.
[0371] [Server operation example] The selection process shown in FIG. 38 can be applied to an operation example of the information processing device 1000. However, in step S1302, it is determined whether or not only image selection information has been received. In addition, in step S1352, a user vector is generated based on the image selection information and text information (or only text information). That is, a user vector C for selecting a new image is generated using a user vector (user vector A) generated based on the image selection information and a user vector (user vector B) generated based on the text information. In this case, the user vector C can be generated using the above-mentioned formula 7.
[0372] In this way, when the user transmits information on his / her preferences (for example, image selection information, text information) to the information processing device 1000, the information processing device 1000 generates a user vector (for example, user vector C) based on each piece of information, and the user can select an image according to his / her preferences based on the user vector and present it to the user. In this case, if there is a favorite image of the user among the images presented to the user, the favorite image is selected and transmitted to the information processing device 1000. In this case, it is possible to generate a user vector that better reflects the user's preferences by adding a new favorite image in addition to the favorite image already selected by the user. In this way, by sequentially adding the favorite images of the user, it is possible to update the user vector so that the user's preferences are better reflected, and it is possible to improve the accuracy of image recommendation. On the other hand, when there is no favorite image of the user among the images presented to the user, if a request (for example, text information) is transmitted to the information processing device 1000 with other information, the information processing device 1000 can recommend an image again reflecting the request. Note that the image selection information also includes selection information when a favorite image of the user is selected from among the images displayed randomly.
[0373] 29, 32, 35, 38, etc. are examples for implementing this embodiment, and the order of some of the steps may be changed within the scope of implementing this embodiment, some of the steps may be omitted, or other steps may be added. In addition, in each of these steps, an example is shown in which the control unit 1150 controls the display state of the display screen displayed on the electronic device 1100 based on information transmitted from the information processing device 1000, but this is not limiting. For example, the display state of the display screen displayed on the electronic device 1100 may be controlled based on the control on the information processing device 1000 side.
[0374] In this way, the user can request the information processing device 1000 to improve the proposed image, generate a desired vector (user vector B) based on the request content, and update the user vector using the desired vector. The updated user vector is user vector C. In this way, by updating the user vector based on the request content from the user, the user can specifically communicate his / her own desires, and can more directly influence the image recommendation results. As a result, it is possible to improve the user's satisfaction and maintain the user's interest.
[0375] [Example of generating tag information using text information converted based on an image] As shown in the first and second embodiments, it is possible to generate text information based on an image, and extract a plurality of feature quantities corresponding to each of a plurality of tags for each tag based on the text information. Therefore, in the third embodiment, an example in which the first and second embodiments are applied is shown here. Note that, among the contents shown below, some of the explanations of the parts common to the first and second embodiments will be omitted.
[0376] For example, when one or more images are selected on the image selection screens shown in FIG. 26, FIG. 36, FIG. 37, FIG. 39, FIG. 40, etc., the tag generating unit 1022 generates text information based on the selected one or more images. For example, when one image is selected, text information corresponding to the image is generated. Also, for example, when multiple images are selected, multiple pieces of text information corresponding to each of the images are generated. Then, the tag generating unit 1022 extracts multiple feature amounts corresponding to each of the multiple tags for each tag based on the generated text information. Note that even when tag information is associated with the selected image in the image information DB 1040, text information may be generated based on the selected image, and tag information may be generated based on the text information (for example, the process of S1303).
[0377] In this case, as in the first and second embodiments, multiple tags may be set based on target information specified by user operation (e.g., a category such as mugs), multiple tags may be set based on the generated text information, or an object (e.g., a mug) contained in the image may be detected and multiple tags may be set based on the detected object.
[0378] Also, similarly to the first and second embodiments, a plurality of pieces of text information may be generated based on one image, and a plurality of feature amounts corresponding to each of a plurality of tags may be extracted for each tag based on the plurality of pieces of text information.
[0379] Then, the vector generating unit 1023 generates a user vector using a plurality of tags generated based on the text information. This user vector is a user vector generated based on the image selection information, and is therefore a user vector A. The method of generating a user vector using a plurality of tags is the same as the above-mentioned generation method. Also, as in the examples shown in Figs. 36 to 40, a user vector C for selecting a new image may be generated using the user vector A and a user vector (user vector B) generated based on text information (e.g., a user's request). Also, the extraction unit 1024 extracts an image or a provider using the user vector (user vector C) generated by the vector generating unit 1023.
[0380] In this way, because the features of multiple tags are extracted for each tag based on the text information generated from the image, it is possible to extract new features that are different from features extracted directly from the image. In other words, it is possible to extract new features that take into account differences caused by differences between languages. In this way, it becomes possible to appropriately extract image features, and it becomes possible to more appropriately select images or providers that suit the user's preferences.
[0381] [Example of setting the reflection ratio using the similarity of user vectors] Also, for example, the similarity between a user vector (user vector A) generated based on image selection information and a user vector (user vector B) generated based on text information may be calculated, and a user vector C may be generated based on the similarity. For example, it is possible to calculate a difference value (for each corresponding component) between each component constituting user vector A and each component constituting user vector B, and calculate this difference value as the similarity. Also, for example, it is possible to calculate the cosine similarity between user vector A and user vector B as the similarity. Note that, as described above, when calculating the cosine similarity, it is preferable to normalize each vector. Also, a high similarity between user vector A and user vector B means that the distance between user vector A and user vector B is close in vector space.
[0382] Then, it is possible to set the reflection ratio α based on the similarity between the user vector A and the user vector B. For example, when the similarity is high (for example, when the similarity is equal to or greater than the first reference value), it means that an image close to the user's request is recommended. Therefore, it is possible to set the reflection ratio of the contents of the user's request text to a value close to 0. On the other hand, when the similarity is low (for example, when the similarity is less than the second reference value (wherein the first reference value>the second reference value)), it means that an image different from the user's request is recommended. Therefore, it is possible to set the reflection ratio of the contents of the user's request text to a value close to 1. Also, when the similarity is moderate (for example, when the similarity is equal to or greater than the second reference value and less than the first reference value), it is possible to set the reflection ratio of the contents of the user's request text within the range of 0 to 1 according to the similarity. In this way, it is possible to obtain the similarity between the user vector B (first feature) and the user vector A (second feature), and obtain the weight (reflection ratio α) between the user vector A and the user vector B based on the similarity. That is, it is possible to check to what extent the recommended image reflects the content of the user's request text, and based on the check result, it is possible to provide the user with a new image again.
[0383] [Example of recommending images using vectors that are not similar to user vectors] In the above, an example is shown in which at least one of the user vectors A to C is used to provide the user with an image similar to the user vector. Here, it is assumed that the user's interest or concern may be directed to something different from the user's initial preference (or image, taste). For example, if a user who feels that modern designs are preferred is shown a classic or retro mug that is the opposite of the user's preference, the user may show a strong interest in the classic or retro mug. Therefore, here, an example is shown in which an image of a vector with a different direction (for example, the opposite direction) from the user vector (at least one of the user vectors A to C) is presented to the user as reference information. For example, it is possible to use a vector D with a low similarity to the user vector (at least one of the user vectors A to C) (for example, a vector whose distance from the user vector (at least one of the user vectors A to C) is equal to or greater than a reference value in the vector space). For example, it is possible to use a vector obtained by converting the user vector (at least one of the user vectors A to C) into a line symmetrical shape (i.e., a vector that is the exact opposite of the user vector) as the vector D, using an average vector (for example, a vector with each component of 0.5) as a reference. Also, for example, a correlation DB that stores information on the correlation between each tag (for example, information indicating the relationship between multiple tags, such as when the value of classic is high, the value of modern is low) may be prepared, and the vector D may be generated using this correlation DB. By using this correlation DB, for example, it is possible to obtain a vector that is the exact opposite of the user vector by calculation. Note that the contents of the correlation DB may be appropriately adjusted according to the preferences of the administrator or user. Also, these examples are only examples, and the vector D may be obtained by other conversion methods. For example, the vector farthest from the user vector in the vector space (or a vector within a predetermined distance from the farthest vector based on the vector farthest from the user vector, a vector that is distant from the user vector by a predetermined distance or more) may be set as the vector D. In these cases, an image (for example, an exact opposite image) that is different in direction from the user vector (at least one of the user vectors A to C) can be selected based on the vector D.That is, an image (for example, an image that is the exact opposite) that is different from the direction of the user vector (at least one of the user vectors A to C) can be presented to the user. In this case, one or more images (first images) selected based on the user vector (at least one of the user vectors A to C) and one or more images (second images) selected as images (for example, an image that is the exact opposite) that are different from the direction of the user vector (at least one of the user vectors A to C) can be displayed so that the user can compare them. For example, a first image display area that displays one or more first images and a second image display area that displays one or more second images can be displayed side by side, either vertically or horizontally. In addition, it is possible to clearly indicate in the first image display area or its vicinity that the image is the user's preference, and to clearly indicate in the second image display area or its vicinity that the image is the image that is different from the direction of the user's preference (for example, an image that is the exact opposite). This allows the user to see an unexpected image, and it is expected that the user will make a new discovery.
[0384] Also, for example, based on the similarity between user vector A and user vector B, it may be determined whether or not to provide the second image to the user. For example, when the similarity is high (for example, when the similarity is equal to or greater than a first reference value), it means that an image determined to be close to the user's preference is provided, but a request sentence is input by the user. For this reason, it is assumed that the user's preference determined by the information processing device 1000 differs from the actual user's preference (request). In such a case, it is possible to provide the user with the first image (the image determined by the information processing device 1000 to be the user's preference) and the second image (the image that is the exact opposite of the first image) to see the user's reaction (the user's actual preference). On the other hand, for example, when the similarity is low (for example, when the similarity is less than a second reference value), it is possible to newly select an image (first image) that the user likes, and provide only the first image to the user to see the user's reaction.
[0385] In addition, when the second image is displayed, it is assumed that the second image is selected as the image of the user's preference. In this case, the weight β3 of the vector D is increased, and a new user vector C (for example, user vector C=β1×user vector A+β2×user vector B+β3×user vector D ... formula 8, where β1 to β3≧0, β1+β2+β3=1) can be generated. The relationship between β1 and β3 can be set, for example, according to the ratio between the number of first images and the number of second images selected by the user after the first and second images are displayed. For example, when the number of first images selected by the user is 3 and the number of second images is 2, it is possible to set β1:β3=3:2. Also, for example, when a new request sentence is not input by the user after the first and second images are displayed, β2 is set to 0. On the other hand, when a new request sentence is input by the user, β1 to β3 can be set based on the new request sentence. In this case, as described above, the reflection ratios β1 to β3 can be set using an AI model (for example, LLM). Note that these weights are merely examples, and weights obtained by other calculations may be used. That is, the information processing device 1000 can include a search unit (for example, the tag generating unit 1022, the vector generating unit 1023, and the extracting unit 1024) that searches for a first image based on a directionality preferred by the user and a second image based on a directionality different from the directionality (for example, the opposite direction, a directionality in a range different from the directionality by a predetermined value or more (for example, an angle different from the directionality by a predetermined value or more)), and a control unit (for example, an information providing unit 1025) that displays the first image and the second image on a display unit (for example, the UI unit 1160 of the electronic device 1100) in a comparable display mode. Also, the above-mentioned first and second images may be searched for within a category (e.g., a category (e.g., mugs) narrowed down in advance based on input text information (e.g., "I want a stylish mug")) set on an initial search screen (e.g., search screen 1200 (see FIG. 25)) and displayed on the display unit, and the first image searched within the category set on the initial search screen and the second image searched outside the category may be displayed on the display unit. The category outside can be set based on each category stored in a preset category DB. For example, a combination of contradictory categories or completely different categories can be set in advance, and the category outside can be set based on this combination. For example, the category "mugs" and the category "clocks" are stored in the category DB as a combination of different categories, and the second image can be searched for within the category "clocks" outside the category "mugs" based on this category DB. In this case, an image in the direction of the above-mentioned user vector (at least one of user vectors A to C) or an image in a different direction (for example, an image that is the exact opposite) can be selected as the second image (one or more images) from within the category "Clock."
[0386] [Example of effect of the third embodiment] In this way, in the third embodiment, by utilizing techniques such as statistical machine learning and analyzing an image selected by a user (e.g., an image of a product), it is possible to provide or recommend to the user an image that the user likes in terms of design. For example, tags related to the design taste of product images are designed, and tags are automatically assigned by learning with a machine learning model for each category. Then, a user can select a favorite image and use the tag assigned to it to quantify the person's preferences (user vector). By using this user vector and calculating the similarity between the image and the provider, it is possible to recommend a new image that is likely to be liked by each person. In other words, it is possible to recommend an image with a similar design taste based on a favorite image (e.g., an image of a product) selected by the user from the perspective of design taste.
[0387] In addition, by utilizing the LLM, it is possible to re-recommend recommended images based on feedback from the user. For example, a user can request improvements to a recommended image (e.g., "I prefer a more modern design") in text form to the LLM, which can then generate a request vector based on the request and use that request to update the user vector. This allows the user to specifically communicate their request, making it possible for that request to have a more direct influence on the recommendation results. As a result, it is possible to improve user satisfaction and sustain interest. In this way, by introducing a text-based feedback mechanism using the LLM into the system, it becomes possible to easily provide feedback on products recommended to users.
[0388] In this way, when a user selects a favorite image, it is possible to recommend images similar to the taste of the selected image and providers corresponding to that taste. In other words, it is possible to recommend images including products or services with a design taste similar to the design taste (design taste of the product or service) of the image selected by the user. This makes it possible to recommend products on an EC site, for example, by taking into account the abstract ideal image and preferences that the user has for the design taste of the product. Furthermore, if the user is not satisfied with the recommended product, the user can feed that back to the system, making it easy to make new recommendations.
[0389] This makes it easier for users to find images that match their design tastes (for example, images of products that match their tastes), which can increase the number of transactions via EC sites. This is considered to be particularly effective in the sale of products that are selected based on design tastes (for example, houses, clothing, furniture, miscellaneous goods, tableware, ornamental plants, paintings, etc.).
[0390] In this way, in the third embodiment, it is possible to recommend products or services by taking into account the abstract ideal image and preferences that the user has regarding the design taste of the product or service. Also, if the user is not satisfied with the product or service recommended to him / her, the user can feed back the recommended product or service to the system to improve the recommendation. In other words, it is possible to appropriately provide information according to the user's preferences.
[0391] [Example of information processing system configuration] In the above, examples have been shown in which tag setting processing, text information generation processing, feature extraction processing, tag generation processing, vector generation processing, extraction processing, etc. are executed in information processing devices 10, 50, 500, 560, 800, 900, 1000, etc., but all or part of each of these processes may be executed in other devices. In this case, an information processing system is configured by each device that executes part of each of these processes. For example, at least part of each process can be executed using various information processing devices such as a server, a device available to a user (e.g., a smartphone, a tablet terminal, a personal computer), and a server connectable via a predetermined network such as the Internet, and various electronic devices.
[0392] Furthermore, a part (or the whole) of an information processing system capable of executing the functions of the information processing devices 10, 50, 500, 560, 800, 900, 1000, etc. may be provided by an application that can be provided via a predetermined network such as the Internet. This application is, for example, SaaS (Software as a Service).
[0393] [Configuration Example of This Embodiment and Its Effects] As described above, the configurations of the information processing devices 10, 50, 500, 560, 800, 900, and 1000 described above can be appropriately combined as necessary in addition to the combinations described above. Therefore, examples taking such combinations into consideration will be described below.
[0394] The information processing system IS1 is an information processing system including an electronic device 1100 used by a user and an information processing device 1000 capable of providing image information to the electronic device 1100. The information processing system IS1 is also an information processing device capable of providing the electronic device 1100 with recommended images (for example, images 1402 to 1404 (see FIG. 39)) searched for using a selection image selected according to the user's preferences. The information processing system IS1 includes an image information DB1040 (an example of a database) that stores images and features of the images in association with each other for each of a plurality of images, an information acquisition unit 1021 (an example of an acquisition unit) that acquires text information indicating a preference for a recommended image (for example, characters in a text information input field 1405 (see FIG. 39)) (or an information acquisition unit 1021 that acquires a selected image selected according to a user's preference (for example, images 1402, 1403 (see FIG. 39)) and text information indicating a preference for the selected image (for example, characters in a text information input field 1405 (see FIG. 39))), and generates a first feature related to the user's preference based on the text information and generates a second feature based on the selected image. The image processing apparatus includes a tag generating unit 1022 that generates a feature (second feature) of the selected image related to what is included in the selected image (e.g., a mug cup) by using the tag, a vector generating unit 1023 (an example of a feature generating unit, and weights (e.g., a reflection ratio α) of each of the first feature and the second feature may be generated based on text information), and an extracting unit 1024 (an example of a selecting unit) that newly selects an image according to the user's preference from among the images stored in the image information DB1040 based on a comparison result obtained by comparing feature information (e.g., a user vector C) indicating the user's preference generated based on the first feature and the second feature (or weights may be used) with the feature of the image (e.g., an image vector). The extracting unit 1024 may select a provider according to the user's preference from among the providers stored in the provider information DB1050 based on a comparison result obtained by comparing feature information (e.g., a user vector C) indicating the user's preference with the feature of the product or service provided by the provider.The second feature is tag information (or a user vector A corresponding thereto) generated based on image information, and the first feature is tag information (or a user vector B corresponding thereto) generated based on text information. A user vector C is calculated by Equation 7 based on the user vector A, the user vector B, and the reflection ratio α. In other words, the vector generating unit 1023 generates feature information (for example, a user vector C) indicating the user's preference based on text information (for example, the first feature) indicating a request for a recommended image and a selected image (for example, the second feature) selected according to the user's preference. In this case, it is possible to generate feature information (for example, a user vector C) indicating the user's preference using the respective weights (for example, the reflection ratio α) of the first feature and the second feature generated based on the text information. That is, the information processing device 1000 is able to newly select an image according to the user's preference based on the text information (for example, the first feature) indicating a request for a recommended image and a selected image (for example, the second feature) selected according to the user's preference. In addition, the second feature may generate text information based on the selected image, and the generated text information may be used as tag information. Moreover, the information processing method according to the present embodiment is an information processing method including each of those processes. Moreover, the program according to the present embodiment is a program that causes a computer to execute each of those processes. In other words, the program according to the present embodiment is a program that causes a computer to realize each function that can be executed by the information processing device 1000.
[0395] According to this configuration, it is possible to select a new image according to the user's preferences using a user vector C generated based on the text information input by the user and the image selected by the user. Therefore, it is possible to propose an appropriate image according to the user's preferences. This makes it possible to appropriately grasp the abstract ideal image and preferences that the user has for the design taste of a product or service, for example, and recommend such a product or service to the user. In addition, it is possible to support the user's various selections in an intuitive and easy-to-understand manner.
[0396] The information processing system IS1 is an information processing system including an electronic device 1100 used by a user and an information processing device 1000 capable of providing the electronic device 1100 with image information. The information processing system IS1 includes an image information DB1040 (an example of a database) that stores images and features of the images in association with each of a plurality of images, an information acquisition unit 1021 (an example of an acquisition unit) that acquires a selected image according to a user's preference from an electronic device 1100, a tag generation unit 1022 (an example of a text information generation unit) that generates text information based on the selected image, the tag generation unit 1022 (an example of a feature generation unit) that generates features corresponding to a predetermined item (one or a plurality of items) based on the text information (for example, extracts a plurality of feature amounts (an example of features) corresponding to each of a plurality of tags (an example of items) for each item) and sets the features of the selected image as features, and an extraction unit 1024 (an example of a selection unit) that newly selects an image according to the user's preference from among images stored in the image information DB1040 based on a comparison result obtained by comparing the features of the selected image (for example, user vector A or C) with the features of the image (for example, vector of the image). The information processing method according to this embodiment is an information processing method including each of these processes. The program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each function that can be executed by the information processing device 1000.
[0397] According to this configuration, it is possible to extract the feature amounts of a plurality of tags for each tag based on the text information generated from the image. For example, a general-purpose AI model can be used for the feature generation unit. In this case, a large amount of teacher data for extracting the feature amounts of the image is not necessary. That is, it is possible to appropriately extract the feature amounts of the image without preparing a large amount of teacher data. In addition, since the feature amounts of a plurality of tags are extracted for each tag based on the text information generated from the image, it is possible to extract new feature amounts different from the feature amounts directly extracted from the image. That is, it is possible to extract new feature amounts taking into account differences caused by differences between languages. In this way, it is possible to appropriately extract the feature amounts of the image. In addition, it is possible to select a new image according to the user's preference from among the plurality of images using the image feature extracted in this way. Therefore, it is possible to propose an appropriate image according to the user's preference. As a result, it is possible to appropriately grasp the abstract ideal image and preference that the user has for the design taste of a product or service, for example, and recommend such a product or service to the user. In addition, it is possible to support various selections of the user in an intuitive and easy-to-understand manner.
[0398] The information processing system IS1 is an information processing system including an electronic device 1100 used by a user and an information processing device 1000 capable of providing image information to the electronic device 1100. The information processing system IS1 includes a provider information DB 1050 (an example of a database) that stores features related to products or services provided by the providers and the providers in association with each of a plurality of providers, an information acquisition unit 1021 (an example of an acquisition unit) that acquires a selected image according to the user's preference from the electronic device 1100, a tag generation unit 1022 that generates a feature of the selected image related to what is included in the selected image (e.g., a mug) based on the selected image, a vector generation unit 1023 (an example of a feature generation unit), and an extraction unit 1024 (an example of a selection unit) that selects a provider according to the user's preference from a plurality of providers stored in the provider information DB 1050 based on a comparison result obtained by comparing the feature of the selected image with the feature related to the product or service. The information processing method according to this embodiment is an information processing method including each of these processes. The program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each function that can be executed by the information processing device 1000.
[0399] According to this configuration, a selected image according to the user's preference is acquired from the electronic device 1100, and a provider according to the user's preference can be selected from among a plurality of providers based on a comparison result obtained by comparing the characteristics of the selected image related to an object included in the acquired image (e.g., a mug) with the characteristics related to the product or service in the provider information DB 1050. Therefore, it is possible to propose a provider appropriate to the user's preference. This makes it possible, for example, to appropriately grasp the abstract ideal image and preferences that the user has regarding the design taste of a product or service, and to recommend such a product or service to the user. Also, it becomes possible to support the user's various selections in an intuitive and easy-to-understand manner.
[0400] The information acquiring unit 1021 (an example of an acquiring unit) acquires a plurality of selected images according to the user's preferences from the electronic device 1100. The information processing system IS1 further includes a vector generating unit 1023 (an example of a feature information generating unit) that generates a user vector (an example of feature information) indicating the user's preferences based on the features of each of the plurality of selected images. The extracting unit 1024 (an example of a selecting unit) newly selects an image according to the user's preferences based on a comparison result obtained by comparing the user vector with the features (e.g., image vectors) of an image in the image information DB 1040 (an example of a database).
[0401] According to this configuration, it is possible to select a new image according to the user's preference from among the multiple images based on the result of comparing a user vector generated based on the characteristics of each of the multiple selected images of the user's preference with the characteristics of the images in the image information DB 1040. Therefore, it is possible to propose a more appropriate image according to the user's preference.
[0402] The information processing system IS1 further includes an information providing unit 1025 that transmits display information for displaying an image newly selected by the extraction unit 1024 (an example of a selection unit) to the electronic device 1100. The electronic device 1100 includes a control unit 1150 that executes display control to cause a UI unit 1160 (an example of a display unit) to display the newly selected image based on the display information.
[0403] According to this configuration, the user can easily visually check, on the UI section 1160 of the electronic device 1100, an image that matches the user's preferences and is newly selected from among a plurality of images based on the user's preferred image.
[0404] The information processing system IS1 further includes an information providing unit 1025 that transmits display information for displaying images stored in an image information DB 1040 (an example of a database) to the electronic device 1100. The information acquiring unit 1021 (an example of an acquiring unit) acquires from the electronic device 1100 an image selected based on a user operation by a user from among the images displayed on the electronic device 1100 based on the display information, as a selected image according to the user's preference.
[0405] According to this configuration, it is possible to display images stored in image information DB 1040 on electronic device 1100. Therefore, the user can easily visually check the images displayed on electronic device 1100, and can easily select a preferred image from among them.
[0406] The electronic device 1100 includes a control unit 1150 that executes display control for displaying images stored in an image information DB 1040 (an example of a database) on a UI unit 1160 (an example of a display unit) based on display information, and transmission control for transmitting edited image information relating to an edited image obtained by editing an image displayed on the UI unit 1160 based on a user operation to the information processing device 1000. Furthermore, an information acquisition unit 1021 (an example of an acquisition unit) acquires an image corresponding to the edited image information transmitted from the electronic device 1100 as a selected image.
[0407] According to this configuration, by enabling editing of the image displayed on electronic device 1100 according to the user's preferences, it is possible to more appropriately express the user's preferences based on the user's subjective editing. Therefore, it is possible to propose appropriate images, etc. according to the user's preferences based on the user's editing results.
[0408] When a re-proposal request is sent from the electronic device 1100 requesting re-proposal of a new image other than the newly selected image, the information providing unit 1025 uses each feature generated by the tag generation unit 1022 and the vector generation unit 1023 (an example of a feature generation unit) to send display information to the electronic device 1100 for displaying the newly selected image as an image according to the user's preferences from among the images stored in the image information DB 1040 (an example of a database).
[0409] According to this configuration, the user can use the electronic device 1100 to execute a re-proposal request to request a re-proposal of a new image other than the image proposed to the user. As a result, even if an image that is not to the user's taste is proposed, it is possible to easily request a re-proposal of an image according to the user's taste. In this case, since the image used for the new selection can be selected based on the characteristics of an image once selected by the user, it is possible to provide the user with an image close to the user's taste. As a result, it is possible to propose a more appropriate image, etc. according to the user's taste.
[0410] The electronic device 1100 includes an image information DB 1080 (an example of an image information database) that stores images acquired by an image acquisition unit 1130, and a control unit 1150 that executes display control for displaying images stored in the image information DB 1080 on a UI unit 1160 (an example of a display unit), and transmission control for transmitting selected image information related to an image selected based on a user operation from among the images displayed on the UI unit 1160 to the information processing device 1000. Also, the information acquisition unit 1021 (an example of an acquisition unit) acquires an image corresponding to the selected image information transmitted from the electronic device 1100 as a selected image.
[0411] According to this configuration, it is possible to more appropriately express the user's preferences based on the user's subjective image collection by allowing the user to select an image according to the user's preferences from among images acquired using the electronic device 1100 (or other external devices). Therefore, it is possible to propose appropriate images, etc. according to the user's preferences based on the user's collection results.
[0412] The tag generator 1022 (an example of a feature generator) may generate at least one of the first feature, the second feature, and the weight using an AI model that extracts features related to what is included in the target image.
[0413] According to this configuration, it is possible to appropriately generate the first feature, the second feature, the weight, and the like for any image according to the user's preference transmitted from the electronic device 1100. Therefore, it is possible to propose an appropriate image, etc. according to the user's preference.
[0414] The information processing device 1000 is capable of providing recommended images searched for using selected images chosen according to the user's preferences to the electronic device 1100 used by the user, and is an information processing device capable of handling an image information DB 1040 (an example of a database) that stores images and their features in association with each of multiple images. The information processing device 1000 includes an information acquisition unit 1021 (an example of an acquisition unit) that acquires text information indicating a request for a recommended image, a tag generation unit 1022 that generates a first feature related to a user's preference based on the text information, generates a feature (second feature) of a selected image related to what is included in the selected image (e.g., a mug) based on the selected image, and generates weights (e.g., a reflection ratio α) for each of the first feature and the second feature based on the text information, a vector generation unit 1023 (an example of a feature generation unit), and an extraction unit 1024 (an example of a selection unit) that newly selects an image according to the user's preference from among images stored in the image information DB 1040 based on a comparison result obtained by comparing feature information (e.g., a user vector C) indicating the user's preference generated based on the weights, the first feature, and the second feature with the feature of the image (e.g., a vector of the image). The first feature is tag information (or a user vector A corresponding thereto) generated based on the image information, and the second feature is tag information (or a user vector B corresponding thereto) generated based on the text information. Moreover, based on the user vector A, the user vector B, and the reflection ratio α, the user vector C is calculated by the formula 7. Moreover, the information processing method according to the present embodiment is an information processing method that causes a computer to execute each of those processes. Moreover, the program according to the present embodiment is a program that causes a computer to execute each of those processes. In other words, the program according to the present embodiment is a program that causes a computer to realize each function that can be executed by the information processing device 1000.
[0415] [Configuration Examples and Effects of the First and Second Embodiments] Conventionally, there are techniques for performing various processes using various information related to images. For example, a technique for extracting multiple feature amounts corresponding to each of multiple images using a trained model has been proposed (for example, JP 2023-000313 A).
[0416] In the conventional technology, it is possible to extract image features using a trained model. For example, in order to generate a trained model for extracting image features of a certain direction, a large amount of training data in which features are associated by human work is required. However, in order to prepare a large amount of training data, a large number of personnel with specialized knowledge are required, and it is generally difficult to obtain a large number of personnel. In addition, if a large amount of training data cannot be prepared, it is expected that it will be difficult to appropriately extract image features.
[0417] Therefore, in this embodiment, the following configuration example makes it possible to appropriately extract image features.
[0418] The information processing devices 50, 500, 560, 800, and 900 are examples of information processing systems that extract features (examples of features) related to tags (examples of items) for things included in an image (e.g., one-piece dresses). For example, the information processing devices 800 and 900 include a tag setting unit 12 (an example of a setting unit) that sets multiple tags based on target information (e.g., a category such as clothing) specified by a user operation, an image acquisition unit 51 that acquires an image, a communication unit 910, a text information generation unit 520 (corresponding to the text information generation unit 520 shown in FIG. 9 and FIG. 15) (an example of a generation unit) that generates multiple pieces of text information related to things included in the image (e.g., one-piece dresses), a feature extraction unit 530 (corresponding to the feature extraction unit 530 shown in FIG. 9 and FIG. 15) (an example of an extraction unit) that extracts multiple features corresponding to each of the multiple tags from the multiple pieces of text information for each tag as a criterion for extracting features of the target information (e.g., from the perspective of a clothing expert), and a DB control unit 810 (an example of a control unit) that associates an image with the multiple features extracted for each tag and outputs them. The information processing devices 800 and 900 shown here may be configured by one device or by multiple devices. Also, instead of the information processing devices 800 and 900, an information processing system configured by multiple devices capable of executing each process realized by the information processing devices 800 and 900 may be used.
[0419] According to this configuration, it is possible to extract the feature amounts of multiple tags for each tag based on the text information generated from the image. For example, it is possible to use a general-purpose AI model for both the text information generating unit and the feature extracting unit. In this case, a large amount of teacher data for extracting the feature amounts of the image is not necessary. That is, it is possible to appropriately extract the feature amounts of the image without preparing a large amount of teacher data. In addition, since the feature amounts of multiple tags are extracted for each tag based on the text information generated from the image, it is possible to extract new feature amounts different from the feature amounts extracted directly from the image. That is, it is possible to extract new feature amounts that take into account differences caused by differences between languages. In addition, it is possible to extract multiple feature amounts from the text information based on target information (e.g., a category such as clothing) specified by a user operation (e.g., from the perspective of a clothing expert), so that it is possible to extract appropriate feature amounts according to the user's preferences. That is, it is possible to appropriately extract the features of the image.
[0420] The information processing device 50, 560, 800, 900 includes an image acquisition unit 51 and a communication unit 910 (an example of an acquisition unit) that acquire an image 41, a text information generation unit 52, 520 (an example of a generation unit) that generates text information 42 based on the image, a feature extraction unit 53, 530 (an example of an extraction unit) that extracts features corresponding to a predetermined item (one or multiple items) based on the text information (for example, extracts multiple feature amounts (an example of features) corresponding to each of multiple tags (an example of items) for each item), and a record control unit 54 and a DB control unit 810 (an example of a control unit) that associate an image with the multiple feature amounts extracted for each item. Note that when one tag is set by the tag setting unit 12 (an example of a setting unit), the feature extraction unit 53, 530 extracts one feature amount corresponding to one tag (an example of a predetermined item) based on the text information. Moreover, the information processing method according to this embodiment is an information processing method including each of these processes. Moreover, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to the present embodiment is a program that causes a computer to realize each function that can be executed by each information processing device.
[0421] According to this configuration, it is possible to extract the feature amounts of a plurality of tags for each tag based on the text information generated from the image. For example, it is possible to use a general-purpose AI model for both the text information generating unit and the feature extracting unit. In this case, a large amount of training data for extracting the feature amounts of the image is not required. That is, it is possible to appropriately extract the feature amounts of the image without preparing a large amount of training data. In addition, since the feature amounts of a plurality of tags are extracted for each tag based on the text information generated from the image, it is possible to extract new feature amounts different from the feature amounts directly extracted from the image. That is, it is possible to extract new feature amounts that take into account differences caused by differences between languages. In this way, it is possible to appropriately extract the feature amounts of the image.
[0422] The information processing device 10, 800 further includes a tag setting unit 12 (an example of a setting unit) that sets multiple tags (an example of an item) based on target information (e.g., a category such as clothing) specified by user operation, and the feature extraction unit 53 may extract multiple features for each tag from the text information 42 as a criterion for extracting features (an example of features) of the target information (e.g., from the perspective of a clothing expert).
[0423] According to this configuration, it is possible to extract multiple features from the text information 42 based on target information (e.g., a category of clothing, etc.) specified by user operation (e.g., from the perspective of a clothing expert), making it possible to extract appropriate features according to the user's preferences.
[0424] The information processing device 800 may further include a tag setting unit 12 (an example of a setting unit) that sets a plurality of tags (an example of an item) based on the text information .
[0425] According to this configuration, it is possible to generate multiple tags based on text information 42 generated from an image 41 specified by a user operation, so that appropriate tags can be set according to the user's preferences and features corresponding to those tags can be extracted.
[0426] The information processing device 800 may further include a tag setting unit 12 (an example of an object detection unit) that detects an object (e.g., a one-piece dress) included in the image 41, and a tag setting unit 12 (an example of a setting unit) that sets multiple tags (an example of items) based on the detected object.
[0427] According to this configuration, it is possible to generate multiple tags based on an object (e.g., a one-piece dress) contained in image 41 specified by user operation, so that appropriate tags can be set according to the user's preferences and features corresponding to those tags can be extracted.
[0428] In addition, the text information generation unit 520 (an example of a generation unit) may generate multiple pieces of text information based on the image 41, and the feature extraction unit 530 (an example of an extraction unit) may extract multiple features (an example of a feature) for each tag (an example of an item) based on the multiple pieces of text information.
[0429] For example, using one text information generation model (e.g., img2txt) heavily depends on the features of that model, whereas using multiple text information generation models (e.g., img2txt) can mitigate this dependency. Also, using multiple models can stabilize behavior.
[0430] The text information generating unit 520 (an example of a generating unit) may generate a plurality of pieces of text information by a plurality of generating processes using a plurality of algorithms.
[0431] According to this configuration, it is possible to generate a plurality of pieces of text information using a plurality of algorithms, and it is possible to stabilize the behavior of the system without depending on the characteristics of each algorithm.
[0432] The feature extraction section 530 (an example of an extraction section) may extract a plurality of features for each tag (an example of an item) based on text information that combines a plurality of pieces of text information.
[0433] According to this configuration, it is possible to extract multiple features for each tag based on text information that combines multiple pieces of text information, making it possible to stabilize the behavior of the system without relying on the features of each piece of text information.
[0434] The feature extraction unit 530 (an example of an extraction unit) may extract a plurality of feature amounts (an example of features) for each tag (an example of an item) for each of the plurality of pieces of text information. In addition, the information processing device 560 may calculate the plurality of feature amounts and weights α i The image processing apparatus may further include a weighted average calculation unit 550 that calculates a plurality of feature amounts for each set of a plurality of tags by executing a predetermined first calculation process (for example, a weighted average calculation process) using the above.
[0435] According to this configuration, the weight α i Since multiple feature amounts for multiple tags can be obtained by weighted average calculation processing using the above, it is possible to improve the calculation accuracy of the feature amounts.
[0436] The text information generating unit 520 may generate the plurality of pieces of text information TX1 to TX4 based on one or a predetermined number of pieces of image information TD2 to which target information (e.g., a category such as clothes) related to a plurality of tags (an example of an item) is set and tag information TD3 (an example of feature information) indicating a plurality of feature amounts (an example of features) related to each of the plurality of tags for an item (e.g., a one-piece dress) included in the image information TD2 is added. Also, the feature extracting unit 530 (an example of an extracting unit) may extract a plurality of feature amounts corresponding to each of the plurality of tags from the plurality of pieces of text information TX1 to TX4 for each tag as a criterion (e.g., from the viewpoint of a clothing expert) when extracting a feature amount of the target information (e.g., a category such as clothes). Also, the information processing device 500 compares the feature amounts of each of the plurality of tags included in the tag information TD3 added to the image information TD2 with the feature amounts of each of the plurality of tags extracted by the feature extracting unit 530 for each tag, and calculates weights α related to the plurality of pieces of text information TX1 to TX4 based on a predetermined second calculation process (a calculation process in which the above-mentioned (1) and (2) are repeated M times) using the comparison result. iThe image forming apparatus may further include a weight calculator 540 for calculating:
[0437] According to this configuration, the weight α used in the weighted average calculation process is calculated using one or a small amount of tagged data TD1. i This makes it possible to appropriately calculate the feature amount, thereby improving the accuracy of the calculation of the feature amount.
[0438] The information processing device 500 includes a text information generating unit 520 (an example of a generating unit) that generates a plurality of pieces of text information TX1 to TX4 based on one or a predetermined number of pieces of image information TD2 to which a plurality of tags (an example of items) related to target information (e.g., a category such as clothing) are set and tag information TD3 (an example of feature information) indicating a plurality of feature amounts (an example of features) related to each of the plurality of tags for an object (e.g., a one-piece dress) included in the image information TD2 is added; a feature extracting unit 530 (an example of an extracting unit) that extracts a plurality of feature amounts corresponding to each of the plurality of tags from the plurality of pieces of text information TX1 to TX4 for each tag as a criterion for extracting a feature amount of the target information (e.g., a category such as clothing) related to the image information TD2 (e.g., from the perspective of a clothing expert); and a feature extracting unit 530 that compares, for each tag, the feature amount for each of the plurality of tags included in the tag information TD3 added to the image information TD2 with the feature amount for each of the plurality of tags extracted by the feature extracting unit 530, and calculates weights α for the plurality of pieces of text information TX1 to TX4 based on a predetermined calculation process (a calculation process in which the above-mentioned (1) and (2) are repeated M times) using the comparison result. i and a weight calculation unit 540 that calculates the weights. Moreover, the information processing method according to this embodiment is an information processing method including each of these processes. Moreover, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each function that can be executed by each information processing device.
[0439] According to this configuration, the weight α used in the weighted average calculation process is calculated using one or a small amount of tagged data TD1. iIn other words, it is possible to appropriately extract image features without preparing a large amount of training data. i Since it is possible to appropriately calculate the feature amount, it is possible to improve the accuracy of calculation of the feature amount. In other words, it is possible to appropriately extract the features of the image.
[0440] Note that each processing procedure shown in this embodiment is an example for realizing this embodiment, and the order of some of the processing procedures may be changed within the scope that allows the realization of this embodiment, and some of the processing procedures may be omitted or other processing procedures may be added.
[0441] Each process shown in this embodiment is executed based on a program for causing a computer to execute each processing procedure. Therefore, this embodiment can also be understood as an embodiment of a program for realizing the function of executing each process, and a recording medium for storing the program. For example, the program can be stored in the storage device of the information processing device by an update process for adding a new function to the information processing device. This makes it possible to cause the updated information processing device to execute each process shown in this embodiment.
[0442] Although the embodiments of the present invention have been described above, the above-mentioned embodiments merely show some of the application examples of the present invention, and it is not intended that the technical scope of the present invention be limited to the specific configurations of the above-mentioned embodiments. [Explanation of symbols]
[0443] 10, 50, 500, 560, 800, 900, 920 Information processing device, 11 Information acquisition unit, 12 Tag setting unit, 13, 54 Recording control unit, 14, 55 Memory unit, 51 Image acquisition unit, 52, 520 Text information generation unit, 53, 530 Feature extraction unit, 100 Tag DB, 200 Image DB, 510 Acquisition unit, 540 Weight calculation unit, 550 Weighted average calculation unit, 600 Weight DB, 810 DB control unit, 820 Output unit, 821, 931 Display unit, 910 Communication unit, NW1 Network, 930 Electronic device, 1000 Information processing device, 1010 Communication unit, 1020 Control unit, 1021 Information acquisition unit, 1022 Tag generation unit, 1023 Vector generation unit, 1024 Extraction unit, 1025 Information providing unit, 1030 storage unit, 1040 image information DB, 1050 provider information DB, 1060 tag information DB, 1100 electronic device, 1110 communication unit, 1120 location information acquisition unit, 1130 image acquisition unit, 1140 sound acquisition unit, 1150 control unit, 1160 UI unit, 1161 reception unit, 1162 output unit, 1170 storage unit, IS1...information processing system
Claims
1. An information processing system including an electronic device used by a user and an information processing apparatus capable of providing a recommended image retrieved using a selection image selected according to the preference of the user to the electronic device, a database that stores, for each of a plurality of images, an association between the image and the characteristics of the image, an acquisition unit that acquires text information indicating a request for the recommended image, a feature generation unit that generates a first feature regarding the preference of the user based on the text information, and generates a second feature that is a feature of the selection image regarding what is included in the selection image based on the selection image, a selection unit that newly selects an image according to the preference of the user from among the images stored in the database based on the feature information indicating the preference of the user generated based on the first feature and the second feature and a comparison result of comparing the feature of the image An information processing system comprising.
2. An information processing system including an electronic device used by a user and an information processing apparatus capable of providing image information to the electronic device, a database that stores, for each of a plurality of images, an association between the image and the characteristics of the image, an acquisition unit that acquires a selection image according to the preference of the user from the electronic device, a text information generation unit that generates text information based on the selection image, a feature generation unit that generates a feature corresponding to a predetermined item based on the text information and uses it as a feature of the selection image, a selection unit that newly selects an image according to the preference of the user from among the images stored in the database based on a comparison result of comparing the feature of the selection image and the feature of the image An information processing system comprising.
3. An information processing system including an electronic device used by a user and an information processing apparatus capable of providing image information to the electronic device, a database that stores, for each of a plurality of providers, an association between the characteristics of a product or service provided by the provider and the provider, an acquisition unit that acquires a selection image according to the preference of the user from the electronic device, a feature generation unit that generates a feature of the selection image regarding what is included in the selection image based on the selection image, a selection unit that selects a provider according to the preference of the user from among the plurality of providers stored in the database based on a comparison result of comparing the feature of the selection image and the feature of the product or service An information processing system comprising
4. The information processing system according to claim 2, wherein the acquisition unit acquires a plurality of the selected images according to the preferences of the user from the electronic device, further comprising a feature information generation unit that generates feature information indicating the preferences of the user based on the features of each of the plurality of selected images, and the selection unit newly selects an image according to the preferences of the user based on a comparison result of comparing the feature information with the features of the image Information processing system.
5. The information processing system according to claim 1 or 2, further comprising an information providing unit that transmits display information for displaying the image newly selected by the selection unit to the electronic device, wherein the electronic device includes a control unit that executes display control for displaying the newly selected image on a display unit based on the display information Information processing system.
6. The information processing system according to claim 1 or 2, further comprising an information providing unit that transmits display information for displaying the image stored in the database to the electronic device, wherein the acquisition unit acquires, as the selected image, an image selected based on a user operation by the user from among the images displayed on the electronic device based on the display information from the electronic device Information processing system.
7. The information processing system according to claim 6, wherein the electronic device includes a control unit that executes display control for displaying the image stored in the database on a display unit based on the display information, and transmission control for transmitting editing image information regarding an edited image obtained by editing the image displayed on the display unit based on a user operation to the information processing apparatus, and the acquisition unit acquires, as the selected image, an image corresponding to the editing image information transmitted from the electronic device Information processing system.
8. The information processing system according to claim 6, wherein when a re-proposal request for re-proposing a new image other than the newly selected image is transmitted from the electronic device, the information providing unit uses each feature generated by the feature generation unit to select, from among the images stored in the database, an image newly selected as an image according to the preferences of the user, and transmits display information for displaying the image to the electronic device Information processing system.
9. The information processing system according to any one of claims 1 to 3, wherein the electronic device An image information database that stores the images acquired by the image acquisition unit, a control unit that executes display control for causing the display unit to display the images stored in the image information database, and transmission control for transmitting, to the information processing apparatus, selection image information regarding an image selected based on a user operation from among the images displayed on the display unit, wherein the acquisition unit acquires, as the selection image, an image corresponding to the selection image information transmitted from the electronic device An information processing system.
10. The information processing system according to claim 1, wherein the feature generation unit generates at least one of the first feature and the second feature using an AI model that extracts features regarding what is included in the target image An information processing system.
11. An information processing system including an electronic device used by a user and an information processing apparatus capable of providing the electronic device with recommended images retrieved using a selection image selected according to the user's preference, a database that stores, for each of a plurality of providers, the features regarding the products or services provided by the providers in association with the providers, an acquisition unit that acquires text information indicating a request for the recommended image, a feature generation unit that generates a first feature regarding the user's preference based on the text information, and generates a second feature that is a feature of the selection image regarding what is included in the selection image based on the selection image, a selection unit that selects a provider according to the user's preference from among the providers stored in the database based on a comparison result of comparing the feature information indicating the user's preference generated based on the first feature and the second feature with the features regarding the product or the service An information processing system comprising.
12. An information processing apparatus capable of providing a recommended image retrieved using a selection image selected according to the user's preference to the electronic device used by the user, and capable of handling a database that stores, for each of a plurality of images, the image and the features of the image in association with each other, an acquisition unit that acquires text information indicating a request for the recommended image, a feature generation unit that generates a first feature regarding the user's preference based on the text information, and generates a second feature that is a feature of the selection image regarding what is included in the selection image based on the selection image, A selection unit that newly selects an image according to the user's preference from among the images stored in the database based on a comparison result of comparing feature information indicating the user's preference generated based on the first feature and the second feature with the features of the image An information processing apparatus comprising the same.
13. An information processing method for causing a computer capable of handling a database that stores, for each of a plurality of images, an image and features of the image and capable of providing a recommended image retrieved using a selected image selected according to a user's preference to an electronic device used by the user, the method comprising: When obtaining text information indicating a request for the recommended image, a feature generation process of generating a first feature regarding the user's preference based on the text information and generating a second feature that is a feature of the selected image regarding what is included in the selected image based on the selected image; A selection process of newly selecting an image according to the user's preference from among the images stored in the database based on a comparison result of comparing feature information indicating the user's preference generated based on the first feature and the second feature with the features of the image An information processing method including the same.
14. A program for causing a computer capable of handling a database that stores, for each of a plurality of images, an image and features of the image and capable of providing a recommended image retrieved using a selected image selected according to a user's preference to an electronic device used by the user, the program comprising: When obtaining text information indicating a request for the recommended image, a feature generation procedure of generating a first feature regarding the user's preference based on the text information and generating a second feature that is a feature of the selected image regarding what is included in the selected image based on the selected image; A selection procedure of newly selecting an image according to the user's preference from among the images stored in the database based on a comparison result of comparing feature information indicating the user's preference generated based on the first feature and the second feature with the features of the image A program for causing a computer to execute the same.