Information processing system, information processing device, information processing method, and program
The information processing system addresses the challenge of personalized product and service recommendations by associating images with features and user preferences, enabling tailored information provision.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BLENDING TECH CO LTD
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies fail to provide information tailored to users' preferences regarding product and service design styles, making it difficult to recommend appropriate products and services.
An information processing system that includes a database associating images with features, an acquisition unit for user preferences, a feature generation unit for generating user and image features, and a selection unit for selecting images based on preference comparisons, enabling personalized image recommendations.
The system effectively provides information tailored to users' preferences, ensuring appropriate product and service recommendations.
Smart Images

Figure 0007861974000005 
Figure 0007861974000006 
Figure 0007861974000007
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing device, an information processing method, and a program for performing image-related processing. [Background technology]
[0002] Conventionally, there are technologies that perform various processes using various information related to images. For example, a technology has been proposed to present multiple images to a user to introduce various products on a website (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2021-082053 [Overview of the project] [Problems that the invention aims to solve]
[0004] For example, in order to understand the abstract ideal image and preferences that users have regarding the design style of products and services, and to recommend such products and services to users, it is important to provide information that is appropriate to the user's preferences.
[0005] The present invention aims to appropriately provide information tailored to the user's preferences. [Means for solving the problem]
[0006] One aspect of the present invention is an information processing system including an electronic device used by a user and an information processing apparatus capable of providing the electronic device with a recommended image retrieved using a selection image selected according to the preference of the user, the information processing system including: a database that stores, for each of a plurality of images, an association between an image and features of the image; an acquisition unit that acquires text information indicating a request for the recommended image; a feature generation unit that generates a first feature regarding the preference of the user based on the text information and generates a second feature that is a feature of the selection image regarding what is included in the selection image based on the selection image; a selection unit that newly selects, from the images stored in the database, an image according to the preference of the user based on feature information indicating the preference of the user generated based on the first feature and the second feature and a comparison result of comparing the feature information with the features of the image; an information processing apparatus constituting the information processing system; an information processing method including each of those processes; and a program causing a computer to execute each of those processes. One aspect of the present invention is an information processing system comprising an electronic device used by a user and an information processing device capable of providing image information to the electronic device, the information processing system comprising: a database that stores images and the features of those images in association with each of a plurality of images; an acquisition unit that acquires selected images from the electronic device according to the user's preferences; a text information generation unit that generates text information based on the selected images; a feature generation unit that generates features corresponding to predetermined items based on the text information and uses them as features of the selected images; and a selection unit that newly selects an image from the images stored in the database according to the user's preferences based on a comparison result obtained by comparing the features of the selected images with the features of the images; an information processing device constituting the information processing system; an information processing method including each of these processes; and a program that causes a computer to execute each of these processes. One aspect of the present invention is an information processing system comprising an electronic device used by a user and an information processing device capable of providing the electronic device with recommended images retrieved using selected images selected according to the user's preferences, the information processing system comprising: a database that stores characteristics of goods or services provided by a provider and the provider associated with each of a plurality of providers; an acquisition unit that acquires text information indicating a request for the recommended images; a feature generation unit that generates a first feature relating to the user's preferences based on the text information and generates a second feature which is a feature of the selected image relating to what is included in the selected image based on the selected image; and a selection unit that selects a provider from among the providers stored in the database that matches the user's preferences based on a comparison result obtained by comparing feature information indicating the user's preferences generated based on the first and second features with the characteristics of the goods or services; an information processing device constituting the information processing system; an information processing method including each of these processes; and a program that causes a computer to execute each of these processes.
Advantages of the Invention
[0007] According to the present invention, information according to the preference of a user can be appropriately provided.
Brief Description of the Drawings
[0008] [Figure 1] It is a block diagram showing a functional configuration example of an information processing apparatus. [Figure 2] It is a block diagram showing a functional configuration example of an information processing apparatus. [Figure 3] It is a diagram showing the stored content of a tag list stored in a tag DB. [Figure 4] It is a diagram showing the stored content of image information and tag information in an image DB. [Figure 5] It is a diagram showing the flow of a setting method for setting a plurality of tags. [Figure 6] It is a diagram showing the flow of an extraction process for extracting feature amounts for each tag. [Figure 7] This is a flowchart illustrating an example of the tag setting process. [Figure 8] This is a flowchart illustrating an example of a feature extraction process. [Figure 9] This is a block diagram showing an example of the functional configuration of an information processing device. [Figure 10] This diagram shows the contents of the weight data stored in the weight database. [Figure 11] This diagram schematically illustrates the flow of the weight calculation process. [Figure 12] This figure shows an example of weight values obtained through weight calculation processing. [Figure 13] This figure shows an example of setting initial values for the weights of a weighted average. [Figure 14] This is a flowchart showing an example of weight calculation processing. [Figure 15] This is a block diagram showing an example of the functional configuration of an information processing device. [Figure 16] This is a flowchart illustrating an example of a feature extraction process. [Figure 17] This is a block diagram showing an example of the functional configuration of an information processing device. [Figure 18] This figure shows an example of how tag information can be displayed on the display unit. [Figure 19] This is a block diagram showing an example of the functional configuration of a communication system. [Figure 20] This figure shows an example of the transitions between display screens shown on the display unit of an electronic device. [Figure 21] This is a block diagram showing an example of the configuration of an information processing system. [Figure 22] This diagram shows the images used for learning and the tags assigned as teacher labels. [Figure 23] This diagram shows a simplified representation of the contents stored in the image information database. [Figure 24] This diagram shows a simplified representation of the contents stored in the provider information database. [Figure 25] This figure shows an example of a search screen displayed on an electronic device. [Figure 26]This figure shows an example of a selection screen displayed on an electronic device. [Figure 27] This figure shows an example of a provider information screen displayed on an electronic device. [Figure 28] This figure shows an example of a selection screen displayed on an electronic device. [Figure 29] This flowchart shows an example of a selection process. [Figure 30] This diagram shows a simplified representation of the contents stored in the item information database. [Figure 31] This figure shows an example of an editing screen displayed on an electronic device. [Figure 32] This flowchart shows an example of image editing processing using electronic devices. [Figure 33] This diagram shows a simplified representation of the contents stored in the image information database. [Figure 34] This figure shows an example of a selection screen displayed on an electronic device. [Figure 35] This flowchart shows an example of image selection processing using electronic devices. [Figure 36] This figure shows an example of a text input screen displayed on an electronic device. [Figure 37] This figure shows an example of a selection screen displayed on an electronic device. [Figure 38] This flowchart shows an example of a selection process. [Figure 39] This figure shows an example of a text input screen displayed on an electronic device. [Figure 40] This figure shows an example of a selection screen displayed on an electronic device. [Modes for carrying out the invention]
[0009] Embodiments of the present invention will be described below with reference to the attached drawings.
[0010] [First Embodiment] [Example of an information processing device configuration] Figure 1 is a block diagram showing an example of the functional configuration of the information processing device 10. The information processing device 10 can be implemented using information processing devices and electronic devices such as servers, personal computers, smartphones, and tablet terminals.
[0011] The information processing device 10 comprises an information acquisition unit 11, a tag setting unit 12, a recording control unit 13, and a storage unit 14. The information acquisition unit 11, the tag setting unit 12, and the recording control unit 13 are each implemented by processing circuits such as one or more CPUs (Central Processing Units) and GPUs (Graphics Processing Units).
[0012] The information acquisition unit 11 is an acquisition unit that acquires character information entered by the user. For example, the information acquisition unit 11 can be an input device capable of inputting character information (e.g., a keyboard or mouse). Alternatively, a voice input device such as a microphone capable of inputting character information by voice, or a voice recognition input device may be used. In addition, an input device consisting of, for example, a camera capable of acquiring character information by capturing images of the character information, or a microphone capable of inputting voice, such as an imaging device, may be used.
[0013] When the information acquisition unit 11 receives character information entered by the user, it outputs that character information to the tag setting unit 12. This character information is used when setting items (tags) that serve as criteria when extracting image features (feature quantities). In this embodiment, as an example of image features, numerically represented features are referred to as feature quantities.
[0014] The tag setting unit 12 generates a list of items to be tagged (tag list) based on the character information output from the information acquisition unit 11, and outputs the character information and information related to the generated tag list to the recording control unit 13. In this embodiment, the items that serve as the basis for extracting image features are referred to as tags. The information associated with tags and their corresponding feature quantities is referred to as tag information. However, tag information can also be referred to as a tag dictionary, metadata, incidental information, additional information, etc. In this embodiment, an example of using numerically represented feature quantities as features corresponding to tags is described. An example of using a score within a predetermined range (0 to 1) as a feature quantity is also shown. The method for setting multiple tags will be explained in detail with reference to Figure 5.
[0015] The recording control unit 13 performs recording control to associate the character information received by the information acquisition unit 11 with the tag list output from the tag setting unit 12 and record it in the tag DB (DataBase) 100. For example, as shown in Figure 3(A), the character information "clothing" received by the information acquisition unit 11 and the tag list "casual, formal, street, business, elegant, vintage, simple, modern, gothic, feminine, bohemian, ethnic, military, rock, punk" set by the tag setting unit 12 are associated and stored in the tag DB 100.
[0016] The memory unit 14 is a storage medium for storing various types of information. For example, the memory unit 14 stores various types of information (e.g., control programs, tag DB 100) necessary for the information acquisition unit 11, tag setting unit 12, and recording control unit 13 to perform various processes. Various storage media can be used as the memory unit 14, such as ROM (Read Only Memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), or combinations thereof.
[0017] Tag DB100 is a database that stores character information entered by the user and a list of tags associated with that information. The list of tags stored in Tag DB100 will be explained in detail with reference to Figure 3.
[0018] Although Figure 1 shows an example in which the information acquisition unit 11 and the storage unit 14 are provided on the information processing device 10, at least one of these may be used as a separate device distinct from the information processing device 10. For example, the information acquisition device and the storage device can be pre-registered with the information processing device 10 (for example, by pairing, connecting via a wired or wireless line) to function as the information acquisition unit and storage unit of the information processing device 10.
[0019] [Example of an information processing device configuration] Figure 2 is a block diagram showing an example of the functional configuration of the information processing device 50. The information processing device 50 can be implemented using information processing devices and electronic devices such as servers, personal computers, smartphones, and tablet terminals.
[0020] The information processing device 50 comprises an image acquisition unit 51, a text information generation unit 52, a feature extraction unit 53, a recording control unit 54, and a storage unit 55. The image acquisition unit 51, text information generation unit 52, feature extraction unit 53, and recording control unit 54 are each implemented by, for example, one or more processing circuits such as CPUs and GPUs. Although Figures 1 and 2 show the information processing device 10 and the information processing device 50 as separate units, they may be configured as an integrated device. An example of such an integrated device is shown in Figure 17.
[0021] The image acquisition unit 51 is an acquisition unit that acquires image information input by the user. For example, an input device capable of receiving image information (e.g., a recording medium reader, a camera) can be used as the image acquisition unit 51. In this embodiment, when referring to an image or image information, it means either the image and the corresponding image file, or either one. For example, it is possible to read and acquire image 41 from a recording medium (e.g., a memory card, USB (Universal Serial Bus) memory, HDD (Hard Disk Drive), CD (Compact Disc), DVD (Digital Versatile Disc), BD (Blu-ray® Disc), etc.) that stores an image file containing a group of multiple images 40 (including image 41). Alternatively, for example, the recording medium containing image 41 and the information processing device 50 can be connected using wireless or wired communication, and the image acquisition unit 51 can read and acquire image 41 from the recording medium. Furthermore, if the recording medium is built into the information processing device 50, the image acquisition unit 51 can read and acquire image 41 from the recording medium. Alternatively, an image 41 may be acquired using an imaging device. For example, the subject included in the image 41 may be captured and acquired using an imaging device. Thus, the image acquisition unit 51 is implemented by an input interface, a file input device, an imaging device, etc.
[0022] The text information generation unit 52 generates text information based on the image information output from the image acquisition unit 51. The text information generation unit 52 then outputs the generated text information to the feature extraction unit 53. This text information generation process will be explained in detail with reference to Figure 6.
[0023] The feature extraction unit 53 extracts multiple feature quantities (feature quantities related to text information) corresponding to multiple tags stored in the tag DB 100 based on the text information output from the text information generation unit 52. The feature extraction unit 53 then outputs the extracted feature quantities of the multiple tags to the recording control unit 54. This process of extracting feature quantities of multiple tags will be explained in detail with reference to Figure 6.
[0024] The recording control unit 54 performs recording control to associate multiple feature quantities for each of the multiple tags extracted by the feature extraction unit 53 with image information acquired by the image acquisition unit 51 and record them in the image DB 200. For example, as shown in Figure 4, the data is stored in the image DB 200.
[0025] The storage unit 55 is a storage medium for storing various types of information. For example, the storage unit 55 stores various types of information (e.g., control program, tag DB 100, image DB 200) necessary for the image acquisition unit 51, text information generation unit 52, feature extraction unit 53, and recording control unit 54 to perform various processing. Various storage media such as ROM, RAM, SRAM, HDD, SSD, or combinations thereof can be used as the storage unit 55. Note that the tag DB 100 corresponds to the tag DB 100 shown in Figure 1.
[0026] In Figure 2, an example is shown in which the image acquisition unit 51 and the storage unit 55 are provided in the information processing device 50. However, at least one of these may be used as a separate device distinct from the information processing device 50. For example, by pre-registering the image acquisition device and the storage device in the information processing device 50, they can function as the image acquisition unit and storage unit of the information processing device 50.
[0027] The image database 200 is a database that stores the feature quantities extracted by the feature extraction unit 53 and the image information acquired by the image acquisition unit 51 in association with each other. The image database 200 will be explained in detail with reference to Figure 4.
[0028] [Example of a tag database configuration] Figure 3(A) is a simplified diagram showing the contents of the target and tag list stored in the tag DB100.
[0029] The tag DB100 is a database for storing multiple tags set by the tag setting unit 12 (see Figure 1).
[0030] Specifically, the target 101 acquired by the information acquisition unit 11 (see Figure 1) and the tag list 102, which shows a list of multiple tags set by the tag setting unit 12, are associated and stored in the tag DB 100.
[0031] As will be explained later, it is possible to generate a list of tags using pre-configured information (for example, a database). An example of the information used in this case is shown in Figure 3(B).
[0032] Figure 3(B) is a simplified diagram showing the contents of the tag information stored in the tag DB100a.
[0033] The tag DB 100a is a database for storing information used when the tag setting unit 12 sets multiple tags (target 101, tag list 102) and information about the multiple tags set by the tag setting unit 12 (selection information 103). The method for setting multiple tags will be explained in detail with reference to Figure 5.
[0034] [Example of image database configuration] Figure 4 is a simplified diagram showing the contents of the image information and tag information stored in the image DB200.
[0035] The image database 200 is a database for storing images acquired by the image acquisition unit 51 and the feature quantities of multiple tags extracted for each tag by the feature extraction unit 53 (see Figure 2), in association with each other.
[0036] Specifically, the image 201 acquired by the image acquisition unit 51, the tag 202 from which feature quantities have been extracted by the feature extraction unit 53, and the multiple feature quantities 203 extracted for each tag are associated and stored in the image DB 200.
[0037] [Example of generating a tag list] Next, we will explain how the tag setting unit 12 sets multiple tags desired by user U1 using the character information entered by user U1.
[0038] Figure 5 shows the flow of a setting method that uses text information entered by user U1 to set multiple tags desired by user U1.
[0039] Figure 5(A) shows an example of user U1 entering the text "clothes". For example, a display screen 16 for entering text information can be displayed on the input / output device 15, and the text "clothes" can be entered using the display screen 16. The input / output device 15 may be provided on the information processing device 10, or an external device different from the information processing device 10 may be used. The input / output device 15 may also utilize a user interface such as a touch panel, or it may be used in conjunction with other operating components (e.g., keyboard, mouse). The subject shown here is used as a criterion when setting tags, and includes meanings such as category and classification.
[0040] When user U1 enters the word "clothes" into the text input field 17 on the display screen 16 and executes the selection operation of the OK button 18, the information acquisition unit 11 receives the text information related to the word "clothes" entered into the text input field 17 and outputs that text information to the tag setting unit 12.
[0041] Figure 5(B) shows a list of tags set by the tag setting unit 12. For example, the tag setting unit 12 can generate a list of tags using an AI (Artificial Intelligent) model (for example, a machine learning model generated through machine learning). In this embodiment, the learning described means finding the regularities behind a large amount of data. Furthermore, the AI model generated by the learning described in this embodiment is generated using various learning algorithms. For example, various images and sentences describing each of these images can be loaded as training data and pre-trained, and this AI model can be used for text information generation processing, etc.
[0042] As an AI model, for example, a large language model (LLM) can be used. Examples of LLMs that can be used include ChatGPT (Generative Pre-trained Transformer), Bard, Llama (Large Language Model Meta AI), Gemini, and Claude. Note that these are just examples, and other AI models may be used. For example, the LLM can be input as instruction information (prompt) that specifies an object (e.g., clothing) and conditions that indicate that it should output multiple tags that can express characteristics related to a specific viewpoint corresponding to that object (e.g., the viewpoint of a clothing expert), and the output information for this can be a list of tags. Here, the specific viewpoint can be defined according to the sensibilities of a person of a certain age with certain attributes (e.g., gender, occupation, background), and can also be called the viewpoint of that person when they look at the object.
[0043] Here, we show an example of a prompt that suggests tags using ChatGPT. For example, the text "clothing" entered in text input field 17 is entered as "category information (e.g., product category)". Then, as an expert on the entered "category information" (e.g., clothing), an instruction (command) is given to display a list of tags that meet the "tag list requirements" for the entered "category information" in the "output format". Here, the "tag list requirements" can be, for example, the words must represent the style or design of the "category information (e.g., product category)".
[0044] Furthermore, the tag setting unit 12 can generate a list of tags using, for example, pre-configured information. For example, as shown in Figure 3(B), it is possible to generate a list of tags using a tag DB 100a that associates the target 101 with the tag list 102. For example, the information acquisition unit 11 can be configured as a reception unit that accepts a selection operation in which the user selects a string corresponding to the target 101 in the tag DB 100a. In this case, the tag setting unit 12 can set the tag list 102 associated with the target 101 corresponding to the string received by the information acquisition unit 11. For example, if "clothing" is received by the information acquisition unit 11, it is possible to set the tag list 102 "casual, formal, ..., rock, punk" associated with the target 101 "clothing". In this case, the selection information 103 in the tag DB 100a is assigned selection information indicating that it has been selected. Figure 3(B) shows an example where "1" is stored as the selection information indicating that it has been selected.
[0045] [Example of generating a tag list using spatial language differences and generational language differences] Here, it is assumed that multiple languages may be used for the characters acquired by the information acquisition unit 11. For example, if the user is Japanese, it is assumed that Japanese, English, and other languages will be entered. Therefore, it is possible to estimate the regionality of the user who performed the input operation based on the entered language. Furthermore, it is assumed that different words may be used depending on the generation for the characters acquired by the information acquisition unit 11. For example, words used only by young people, words understood only by elderly people, etc. If these words are entered, it is possible to estimate the user's generation. For example, if the Japanese word "AB" (where AB is Japanese), which is assumed to be frequently used by women in their 20s, is entered, it is possible to estimate the user's attributes as "Japanese," "in their 20s," and "female." Also, for example, if the Japanese word "CD" (where CD is Japanese), which is assumed to be frequently used by men in their 70s, is entered, it is possible to estimate the user's attributes as "Japanese," "elderly," and "male." In this way, it is possible to estimate the user's attributes from the characters acquired by the information acquisition unit 11, so tags may be set considering the user's attributes. For example, using LLM, it is possible to take words used by a user as input and output user attributes. Therefore, using LLM, it is possible to set tags that take into account user attributes that can be estimated based on the words used by the user. This makes it possible to set appropriate tags that take into account spatial and generational language differences.
[0046] Furthermore, the tag setting unit 12 outputs tag list information for the multiple tags that have been set to the recording control unit 13. Then, as shown in Figure 5(B), the recording control unit 13 stores the multiple tags corresponding to the tag list information output from the tag setting unit 12 in the tag DB 100. In Figure 5, an example is shown where the tags "casual, formal, ..., rock, punk" are determined to correspond to the text "clothes" entered by the user. These contents correspond to Figure 3(A).
[0047] Figure 5 shows an example of setting multiple tags based on the text "clothes" entered by user U1, but other setting methods may be used. For example, the tag setting unit 12 determines other categories related to "clothes" (for example, categories that consider the era, region, and appearance) based on the text information "clothes" output from the information acquisition unit 11. For example, a category DB can be created in advance to store a large amount of text information (for example, clothes, hats, shoes, bags, socks), the categories to which these text information belong (for example, clothing, accessories, fashion), and multiple tags corresponding to the categories (for example, "casual, formal, ..., rock, punk"), and the category can be determined based on this category DB. Alternatively, other AI technologies may be used to determine the category related to the entered text information. The tag setting unit 12 can then extract multiple tags corresponding to the determined category from the category DB and create a tag list from these multiple tags.
[0048] [Example of feature extraction for each tag] Next, we will explain an extraction method in which the information processing device 50 extracts feature quantities for each tag associated with the image information input by user U1.
[0049] Figure 6 shows the flow of the extraction process for extracting feature quantities for each tag related to image 41.
[0050] Figure 6(A) shows image 41 (the image to be tagged) input by user U1. Image 41 is acquired by image acquisition unit 51.
[0051] Figure 6(B) shows the text information 42 generated by the text information generation unit 52 based on the image 41. When the image 41 acquired by the image acquisition unit 51 is input to the text information generation unit 52, the text information generation unit 52 generates text information 42 based on the image 41. The text information shown here is, for example, a sentence formed by characters such as words and sentences written in a predetermined language that humans can handle. Specifically, the text information generation unit 52 generates text as text information 42 for objects included in the image 41 (for example, the subject and its background in the case of an captured image, or the model image (or motif image) and its background image in the case of a generated image, etc.), describing each object itself (for example, appearance, form, color, movement, impression, attributes (for example, properties and characteristics belonging to the object itself)), describing its relationship with other objects, describing events that occur as a result of the interaction of each of these objects, etc. In other words, the text information generation unit 52 generates text as text information 42 that describes one or more objects included in the image 41. Furthermore, it is possible to use a known text generation method for generating this text information.
[0052] For example, the text information generation unit 52 can generate text information 42 based on the image 41 using a pre-prepared database. For example, the text information generation unit 52 can detect one or more objects (and their attributes, size, position in the image, etc.) contained in the image 41, and generate text by connecting descriptions of those objects according to predetermined rules based on this information.
[0053] Furthermore, for example, the text information generation unit 52 can generate text information 42 based on the image 41 using an AI model. Examples of AI models that can be used include image to text and img2txt. These can be implemented by installing a predetermined application on the information processing device 50.
[0054] Furthermore, it is possible to generate text information using other AI models. For example, a multimodal LLM can be used. A multimodal LLM consists of an image processing part (a pre-trained part) used for image recognition processing, a large-scale language model part, and an adapter part that connects them. When the image 41 acquired by the image acquisition unit 51 is input to the multimodal LLM, it is possible to output text information related to the image 41. When using an LLM such as a multimodal LLM, the amount of computation is large, so it is expected that the processing load will be large. Therefore, for example, if the information processing device 50 is a device with a slow processing speed such as a personal computer, a device with a fast processing speed (for example, a server) may be used. For example, the text information generation unit 52 can send the image 41 acquired by the image acquisition unit 51 to the server, the server will perform text information generation processing, and the text information generation unit 52 can receive the processing result (text information) from the server and use it.
[0055] Furthermore, text information may be generated using an object specified by the user, multiple tags set by the tag setting unit 12, etc. For example, when using a multimodal LLM, it is possible to generate text information considering an object specified by the user (e.g., clothing). For example, as text information related to an image 41 acquired by the image acquisition unit 51, conditional information indicating that the LLM should output a text that can express a specific viewpoint corresponding to an object specified by the user (e.g., clothing) can be input to the LLM as instructional information (prompt), and the output information for this can be used as text information. Also, for example, it is possible to generate text information considering multiple tags set by the tag setting unit 12 (e.g., casual, ..., punk). For example, instructional information (prompt) indicating that the LLM should generate a text related to multiple tags (e.g., casual, ..., punk) can be input to the multimodal LLM, and the output information for this can be used as text information. Furthermore, by inputting both of these instructional information (prompts) to the multimodal LLM and using the output information for this as text information, it is possible to generate text information that considers both an object specified by the user (e.g., clothing) and multiple tags (e.g., casual, ..., punk).
[0056] For example, when a person looks at an object (e.g., image 41), the superficial emotions associated with that object (e.g., cute, fun, sad) are easy to put into words. In other words, superficial emotions are easy for humans to express in words. On the other hand, complex emotions based on various memories born from a person's past experiences (e.g., intrinsic emotions, subtle emotions) are often difficult to put into words. It is possible to express such intrinsic emotions and emotions that humans cannot put into words by writing text based on images, and these expressed emotions can be used as the target of analysis for extracting features from multiple tags. That is, when extracting features from multiple tags related to an image, by using text information written from the perspective of a certain human viewpoint, emotion, image, etc., it becomes possible to extract each feature based on the limited content related to the image. This makes it possible to extract appropriate features from a given viewpoint for multiple tags corresponding to a particular viewpoint.
[0057] Figure 6(C) shows the tag information 43 generated based on the feature extraction unit 53. When the text information 42 generated by the text information generation unit 52 is input to the feature extraction unit 53, the feature extraction unit 53 extracts feature quantities for each of the multiple tags stored in the tag DB 100 based on the text information 42. The feature quantities shown here represent a predetermined range of values (scores) that indicate the relationship (or relationship, correlation) between the tags and what is included in the text information 42 (for example, one or more objects, their backgrounds). In this embodiment, an example is shown in which 0 to 1 is used as the predetermined range of values. Furthermore, the value becomes larger as the relationship with the tag increases (maximum value 1), and smaller as the relationship with the tag decreases (minimum value 0). For example, the numerical value "0.4" for the tag "casual" indicates a numerical value related to "casual," and the numerical value "0.7" for the tag "formal" indicates a numerical value related to "formal." For example, since the numerical value "0.7" for the tag "formal" is a relatively high value, it is estimated that the relationship between the object (dress) included in the text information 42 and "formal" is high. On the other hand, for example, the value "0.0" for the tag "rock" indicates a value related to "rock," and the value "0.0" for the tag "punk" indicates a value related to "punk." Since these values are 0, it is presumed that there is little to no association between the object (one-piece dress) contained in the text information 42 and "rock" or "punk." It should be noted that known feature extraction methods can be used for this feature extraction.
[0058] For example, the feature extraction unit 53 can extract features for multiple tags based on the text information 42 using a pre-prepared database. For example, the feature extraction unit 53 uses the database to detect one or more strings related to the character examples of each tag from among the strings contained in the text information 42. The feature extraction unit 53 can then calculate a score for each tag based on the number of strings detected for each tag, the content of the strings, etc. For example, a tag with a large number of detected strings can be calculated with a higher score corresponding to that number. Also, for example, a tag with important content in the detected strings can be calculated with a higher score corresponding to its importance.
[0059] Furthermore, it is possible to extract feature quantities for each tag based on text information 42, for example, by using artificial intelligence (AI). For example, it is possible to extract feature quantities for each tag using an AI model. As this AI model, for example, the aforementioned LLM can be used.
[0060] For example, the feature extraction unit 53 can use an LLM to extract features for multiple tags based on the text information 42. As mentioned above, examples of LLMs that can be used include ChatGPT, Bard, Llama, Gemini, and Claude. These are just examples, and other models may be used. For example, it is possible to input instruction information (prompt) to the LLM, using the sentences that make up the text information 42, to output a score (0 to 1) regarding the relevance of each of the multiple tags, and to use the output information as features for each of the multiple tags. Alternatively, for example, it is possible to input condition information to the LLM as instruction information (prompt), using the sentences that make up the text information 42, to output a score (0 to 1) regarding the relevance of each of the multiple tags from a specific perspective (e.g., the perspective of a clothing expert) corresponding to an object specified by the user (e.g., clothing), and to use the output information as features for each of the multiple tags.
[0061] When using LLM, the computational load is large, so it is expected that the processing load will be high. Therefore, if the information processing device 50 is a device with a slow processing speed, such as a personal computer, a device with a fast processing speed (e.g., a server) may be used. For example, the feature extraction unit 53 can send the text information 42 output from the text information generation unit 52 to the server, where the server performs feature extraction processing for each of the multiple tags, and the feature extraction unit 53 can receive the processing results (feature quantities for each of the multiple tags) from the server and use them.
[0062] Here, we show an example of a prompt that tags text using ChatGPT. For example, the text "clothing" entered in the text input field 17 is entered as "category information," and the text information 42 generated by the text information generation unit 52 is entered as "text information." Furthermore, as an expert in the entered "category information" (e.g., clothing), an instruction (command) is given to receive the entered "text information," convert the content of the entered "text information" into a "vector" that represents the style and atmosphere of the entered "category information," and display that "vector" in the "output format." Here, as a definition of that "vector," it is possible to give an instruction (command) that each element of the "vector" is a real number between 0 and 1, and that the "vector" has elements of multiple tags set by the tag setting unit 12, such as "casual, ..., punk."
[0063] Figure 6(D) shows the relationship between the image 41 acquired by the image acquisition unit 51 and the tag information 43 generated by the feature extraction unit 53. When the tag information 43 generated by the feature extraction unit 53 is input to the recording control unit 54, the recording control unit 54 associates the image 41 output from the image acquisition unit 51 with the tag information 43 and stores it in the image DB 200. For example, as shown in Figure 4, the image 41 and the tag information 43 are stored in association. Although Figure 6 shows an example of associating the image 41 with the tag information 43 and storing it in the image DB 200, the image 41 and the tag information 43 may also be associated and output from an output device or output to another device. For example, as shown in Figures 18 and 20, it is possible to associate the image 41 with the tag information 43 and display it on the display units 821 and 931. This makes it possible for user U1 to easily grasp the feature quantities of multiple tags related to the target specified by user U1 for the image 41. Although the above example demonstrates setting multiple tags and extracting features for each of them, this embodiment can also be applied to setting a single tag and extracting features for that tag.
[0064] [Example of information processing device operation] Figure 7 is a flowchart illustrating an example of the tag setting process in the information processing device 10. This tag setting process is executed based on a program stored in the storage unit 14. This tag setting process is also executed when a user operation is received by the information acquisition unit 11. This tag setting process will be explained with reference to Figures 1 to 6 as appropriate.
[0065] In step S401, the information acquisition unit 11 receives the characters entered by user U1 and outputs the character information related to the received characters to the tag setting unit 12.
[0066] In step S402, the tag setting unit 12 sets a tag list based on the character information output from the information acquisition unit 11 and outputs the set tag list to the recording control unit 13. This setting method is the same as the example shown in Figure 5.
[0067] In step S403, the recording control unit 13 stores the tag list set by the tag setting unit 12 in the tag DB 100. For example, as shown in Figure 3(A), the tag list is stored in the tag DB 100.
[0068] [Example of information processing device operation] Figure 8 is a flowchart illustrating an example of feature extraction processing in the information processing device 50. This feature extraction processing is executed based on a program stored in the storage unit 55. This feature extraction processing is also executed when an image is acquired by the image acquisition unit 51. This feature extraction processing will be explained with reference to Figures 1 to 7 as appropriate.
[0069] In step S411, the image acquisition unit 51 acquires the image input by user U1 and outputs image information related to the acquired image to the text information generation unit 52.
[0070] In step S412, the text information generation unit 52 generates text information based on the image information output from the image acquisition unit 51 and outputs the generated text information to the feature extraction unit 53. The method for generating this text information is the same as in the example shown in Figure 6.
[0071] In step S413, the feature extraction unit 53 extracts feature quantities for each of the multiple tags stored in the tag DB 100 based on the text information generated by the text information generation unit 52. The feature extraction unit 53 then outputs the extracted feature quantities for each of the multiple tags to the recording control unit 54. The method for extracting feature quantities for each tag is the same as in the example shown in Figure 6.
[0072] In step S414, the recording control unit 54 stores the feature quantities for each of the multiple tags output from the feature extraction unit 53 and the image information output from the image acquisition unit 51 in the image DB 200. For example, as shown in Figure 4, the feature quantities for each of the multiple tags and the image information are associated and stored in the image DB 200.
[0073] [Example of generating text information in multiple languages] Here, when representing the same image, the representation often differs depending on the language, and the impression that humans perceive also differs. For example, when representing the same image in Japanese and English, the representation will differ in Japanese and English, and the impression that humans perceive will also differ. Furthermore, it is conceivable that the text information generation unit 52 may be capable of generating text information in multiple languages. For example, it is conceivable that text information could be generated in each language such as Japanese, English, French, and Spanish. In this case, the text information generation unit 52 may generate text information in multiple languages based on the image 41, and the feature extraction unit 53 may extract feature quantities for each of the multiple tags stored in the tag DB 100 based on the text information in multiple languages generated by the text information generation unit 52. In this case, for example, instruction information (prompt) to output scores (feature quantities) for multiple tags using each sentence that constitutes the text information in multiple languages can be input to the LLM, and the output information in response to this can be used as feature quantities for each of the multiple tags.
[0074] [Example of setting tags based on an image] The above example demonstrates setting multiple tags based on an object specified by user U1 and extracting features related to those tags. However, it is also possible to set multiple tags based on some information about the image and extract features related to those tags. For example, object detection processing may be performed on image 41 acquired by image acquisition unit 51, and tags may be set based on the detected object. For example, since image 41 contains a dress, the dress is detected from image 41 by the object detection processing. Furthermore, the attributes of the detected dress can be detected by image recognition processing. For example, the color, shape, size, etc., of the dress can be detected. Therefore, it is possible to determine what kind of clothing the detected dress is. For example, if image 41 contains a long-sleeved green dress with an A-line silhouette and a round neck, it can be estimated that a tag related to "clothing" should be set for image 41.
[0075] Furthermore, if an image contains, for example, a sofa, chairs, and a table, these can be detected from the image through object detection processing. Additionally, the attributes of the detected objects (e.g., size, color, and position) can be detected through image recognition processing. Thus, if an image contains a sofa, chairs, and a table, it can be inferred that a tag related to "furniture" should be set for the image.
[0076] Furthermore, for example, if an image contains a house, the object detection process can detect the house from the image. Additionally, the image recognition process can detect the attributes of the detected house (e.g., size, color, location). Thus, if an image contains a house, it can be assumed that a tag related to "house" should be set for the image. This tag setting based on object detection can be achieved by including the object detection unit in the tag setting unit 12 (see Figure 17).
[0077] [Example of setting tags based on text information] This section demonstrates an example of setting multiple tags based on text information generated from an image and extracting features related to those tags. For example, character recognition processing may be performed on the text information 42 generated by the text information generation unit 52, and tags may be set based on the strings, words, etc., contained in the text information 42. For example, since image 41 contains a dress, it is presumed that the text information 42 also contains a description of the dress. It is also presumed that the text information 42 contains a description of the attributes of the dress. In this case, the character recognition processing can detect the description of the dress, the description of the attributes of the dress, etc., contained in the text information 42. Based on this information, multiple tags can be set, similar to when setting tags based on an image. For example, if the text information 42 contains a description of a green, long-sleeved dress with an A-line silhouette and a round neck, it is possible to set a tag related to "clothing" for image 41.
[0078] [Example of extracting features using a single multimodal LLM] Here, it is also possible to have a single multimodal LLM handle both the text information generation process by the text information generation unit 52 and the feature extraction process by the feature extraction unit 53. In this case, an image may be input to the multimodal LLM, and prompts for generating feature quantities for each of the multiple tags from the image may be input to extract feature quantities related to the image. In this case, both the text information generation process and the feature extraction process are executed in the multimodal LLM, and feature quantities related to the image are extracted based on the results of each of these processes. In other words, the multimodal LLM can function as both the text information generation unit 52 and the feature extraction unit 53.
[0079] When using a multimodal LLM, it is possible to input prompts that simultaneously perform text generation and feature extraction for each tag. For example, the text "clothing" entered in the text input field 17 can be input as "category information," and the image 41 acquired by the image acquisition unit 51 can be input as the "input image." Furthermore, as an expert in the input "category information" (e.g., clothing), an instruction (command) can be given to receive the input "input image," convert the objects shown in the input "input image" into a "vector" that represents the style and atmosphere of the input "category information," and then display that "vector" in the "output format." Here, as a definition of that "vector," it is possible to give an instruction (command) that each element of the "vector" is a real number between 0 and 1, and that the "vector" has elements of multiple tags set by the tag setting unit 12, such as "casual, ..., punk."
[0080] [Example of the effects of the first embodiment] As described above, according to this embodiment, when extracting features related to multiple tags for a subject contained in an image, it is possible to extract the features of multiple tags for each tag based on text information generated from the image. That is, it is possible to extract the features of multiple tags related to an image for each tag by using a text information generation unit 52 that generates text information based on an image and a feature extraction unit 53 that extracts the features of multiple tags for each tag based on the text information. For example, it is possible to use a general-purpose AI model for both the text information generation unit 52 and the feature extraction unit 53. In this case, a large amount of training data for extracting image features is not required. That is, it is possible to appropriately extract image features without preparing a large amount of training data. Furthermore, since the features of multiple tags are extracted for each tag based on text information generated from the image, it is possible to extract new features that are different from the features extracted directly from the image. That is, it is possible to extract new features that take into account differences arising from differences between languages, etc. Thus, according to the first embodiment, it is possible to appropriately extract the features of an image.
[0081] Furthermore, by using technologies such as LLM and multimodal LLM, it is possible to improve the processing speed of feature extraction from multiple images. In other words, it is possible to make the feature extraction process for multiple images more efficient. Also, by using technologies such as LLM and multimodal LLM, it is possible to significantly reduce the amount of manual work required. In other words, it is possible to achieve significant labor savings.
[0082] [Second Embodiment] In the first embodiment, an example was shown in which tag information is generated using one type of text information generation unit 52 and one type of feature extraction unit 53. However, tag information may also be generated using two or more types of text information generation units or two or more types of feature extraction units. Alternatively, tag information may be generated using both two or more types of text information generation units and two or more types of feature extraction units. Therefore, Figure 9 and others show an example in which tag information is generated using two or more types of text information generation units and one type of feature extraction unit.
[0083] [Example of an information processing device configuration] Figure 9 is a block diagram showing an example of the functional configuration of the information processing device 500. Note that the information processing device 500 is a modified version of the information processing device 50 shown in Figure 2. Except for the omission of the image DB 200, the replacement of the image acquisition unit 51, text information generation unit 520, and feature extraction unit 530 with an acquisition unit 510, a text information generation unit 520, and a feature extraction unit 530, and the addition of a weight calculation unit 540 and a weight DB 600, it is identical to the information processing device 50. Therefore, parts common to the information processing device 50 are denoted by the same reference numerals as those in the information processing device 50, and some of their descriptions are omitted.
[0084] The information processing device 500 comprises an acquisition unit 510, a text information generation unit 520, a feature extraction unit 530, a weight calculation unit 540, a recording control unit 54, and a storage unit 55. Each of the acquisition unit 510, text information generation unit 520, feature extraction unit 530, weight calculation unit 540, and recording control unit 54 is implemented by, for example, one or more processing circuits such as CPUs and GPUs. Also, although the information processing device 10 and the information processing device 500 are shown separately, as in the example shown in Figure 2, they may be configured as a single integrated device.
[0085] The acquisition unit 510 is an acquisition unit that acquires tagged data TD1 to which image information TD2 and tag information TD3 are associated. The acquisition unit 510 then outputs image information TD2 to the text information generation unit 520 and tag information TD3 to the weight calculation unit 540. For example, the acquisition unit 510 can acquire an image file to which image information TD2 and tag information TD3 are associated. Note that tagged data TD1 is data used when determining weights (weight values), and is equivalent to, for example, training data for a neural network. However, in the second embodiment, an example is shown in which only 1 or a small amount (less than a predetermined number) of training data is used. For example, the number of tagged data TD1 can be several times to tens of times the number of generation units (i.e., the number of weights (number of models)) that make up the text information generation unit 520.
[0086] The text information generation unit 520 comprises multiple generation units (first generation unit 521 to the Nth generation unit 524). Each of these generation units (first generation unit 521 to the Nth generation unit 524) generates text information based on the image information TD2 output from the acquisition unit 510. The text information generation unit 520 then outputs the text information generated by each generation unit to the feature extraction unit 530. Hereinafter, N is an integer of 2 or more. That is, although Figure 9 shows an example of 4 generation units (first generation unit 521 to the Nth generation unit 524), it may also have 2, 3, or 5 or more generation units. Furthermore, the first generation unit 521 to the Nth generation unit 524 execute processing to convert the image into text information using different algorithms. In this way, when generating text information based on the same image, multiple pieces of text information are generated using different algorithms, each resulting in a different sentence. By using the text information of these different sentences to extract feature quantities for each tag, it becomes possible to extract content that cannot be extracted from a single piece of text information.
[0087] As mentioned above, it is also possible to generate text information in multiple languages based on a single image. Therefore, the first generation unit 521 to the Nth generation unit 524 may generate text information in multiple languages based on a single image. In this way, by generating text information in multiple languages based on the same image, multiple pieces of text information are generated, each containing a different sentence.
[0088] As described above, in the first embodiment, since a single model of text information generation unit 52 (e.g., img2txt) is used, it is considered that there is a great deal of dependence on the characteristics of that model. In contrast, in the second embodiment, multiple models of text information generation units (e.g., img2txt) are used. This makes it possible to mitigate the dependence on the characteristics of a single model. Furthermore, by using multiple models in this way, it is possible to stabilize the behavior. The text information generation process using multiple generation units will be explained in detail with reference to Figure 11, etc.
[0089] The feature extraction unit 530 extracts feature quantities for each tag based on the tag list stored in the tag DB 100 and the multiple text information output from the text information generation unit 520. The feature extraction unit 530 then outputs each extracted feature quantity to the weight calculation unit 540. This feature extraction process will be explained in detail with reference to Figure 11, etc.
[0090] The weight calculation unit 540 calculates the weights used by the weighted average calculation unit 550 (see Figure 15) for each of the multiple generation units (first generation unit 521 to the nth generation unit 524) based on the feature quantities for each tag related to the multiple text information output from the feature extraction unit 530 and the tag information TD3 output from the acquisition unit 510. The weight calculation unit 540 then outputs the calculated weight values to the recording control unit 54. The weight calculation process will be explained in detail with reference to Figure 13, etc.
[0091] The memory unit 55 stores various information (for example, control programs, tag DB100, weight DB600) necessary for the acquisition unit 510, text information generation unit 520, feature extraction unit 530, weight calculation unit 540, and recording control unit 54 to perform various processes.
[0092] The weight DB600 is a database that stores the weight values obtained by the weight calculation unit 540 in association with multiple generation units (first generation unit 521 to the nth generation unit 524). The weight DB600 will be explained in detail with reference to Figure 10.
[0093] [Example of a weight database configuration] Figure 10 is a simplified diagram showing the contents of the weight data stored in the weight DB600. The weight DB600 is a database for storing weight values used when the weighted average calculation process is performed by the weighted average calculation unit 550 (see Figure 15).
[0094] Specifically, the generation unit identification information 601 and the weight 602 are associated and stored in the weight DB 600.
[0095] The generation unit identification information 601 stores information for identifying the multiple generation units (first generation unit 521 to the nth generation unit 524) that make up the text information generation unit 520. Note that in Figure 10, for the sake of simplicity, only the names corresponding to each generation unit are shown.
[0096] Weight 602 stores the weight values obtained by the weight calculation unit 540. Here, the sum of each weight (α1 + α2 + α3 + ... + α N An example is shown where ) is 1. Also, weight α i The specific calculation method is shown in Figures 11 and 13. Note that i is an integer satisfying 0 ≤ i ≤ N.
[0097] [Example of weight calculation process] Figure 11 schematically shows the flow of weight calculation processing by the information processing device 500. Figure 12 shows an example of weight values obtained by weight calculation processing by the information processing device 500.
[0098] As shown in Figure 11, when image information TD2, which constitutes tagged data TD1, is input to the text information generation unit 520, the first generation unit 521 to the Nth generation unit 524 generate text information TX1 to TX4 based on the image information TD2. Note that in Figure 11, for the sake of clarity, only the text related to text information TX1 and TX2 is shown in a simplified form, and the text related to text information TX3 and TX4 is omitted.
[0099] Next, when the text information TX1 to TX4 generated by the first generation unit 521 to the nth generation unit 524 is input to the feature extraction unit 530, the feature extraction unit 530 generates tag information FQ1 to FQ4 based on the text information TX1 to TX4. In Figure 11, for the sake of clarity, the tag information FQ1 to FQ4 is shown in a simplified form within the rectangle representing the feature extraction unit 530. Furthermore, only the data related to tag information FQ1 and FQ2 is shown in a simplified form, while the data related to tag information FQ3 and FQ4 is omitted.
[0100] Next, the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the tag information TD3 that constitutes the tagged data TD1 are input to the weight calculation unit 540.
[0101] Next, the weight calculation unit 540 calculates the weight values for each generation unit (first generation unit 521 to the nth generation unit 524) (weight values for each pipeline) based on the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the tag information TD3 that constitutes the tagged data TD1. Here, we will explain the method for determining the weighted average weights from the loss function. For example, similar to the method used in deep learning, the weighted average weights can be determined using gradient descent.
[0102] Here, the value obtained by extracting only the feature amount of the tag information is referred to as a tag value vector and will be described. Specifically, the tag value vector is a vector whose dimensionality is the number of tags set by the tag setting unit 12. For example, in FIGS. 9 to 11, an example in which the number of tags is 15 is shown, so the tag value vector is a 15-dimensional vector. Therefore, hereinafter, an example in which the tag value vector is a 15-dimensional vector will be shown.
[0103] Also, the tag value vectors (output values of the feature extraction unit 530) generated for each generation unit (the first generation unit 521 to the Nth generation unit 524) are vector x i (x i1 , x i2 , …, x i15 ), and the weight of vector x i is α i . However, i is an integer satisfying 0 ≦ i ≦ N. Note that vector x i can also be referred to as the tag value vector of the output of each pipeline (the first generation unit 521 to the Nth generation unit 524 → the feature extraction unit 530). Also, the tag value vector output from the feature extraction unit 530 is vector y (y1, y2, …, y 15 ). That is, the following Equation 1 holds. Note that the actual vector x i and vector y are column vectors, but here, for ease of explanation, vector x i and vector y are shown in the form of row vectors.
[0104]
Equation
[0105] Also, the tag value vector of the tagged data (corresponding to the teacher data) is vector t (t1, t2, …, t 15 ).
[0106] Next, the procedure of the determination method for determining the weight of the weighted average will be described.
[0107] First, the weight α iDetermine the initial value of . For example, it can be determined based on Equation 2 below. Note that N represents the number of pipelines. α i =1 / N Equation 2
[0108] Next, we define the loss function L. Here, the loss function L is defined based on the following equation 3. L = |vector t - vector y| 2 formula 3
[0109] Next, each α in the loss function L i The derivative with respect to is obtained using the following equation 4.
[0110]
number
[0111] Next, we will use the tagged data (equivalent to training data) as a reference for each α. i Update the data as shown in (1) and (2) below. Repeat this update for each tagged data (equivalent to training data). This single operation is counted as one epoch.
[0112] (1) Updating each value: Update according to the following equation 5. Here, ε is the learning rate (hyperparameter). This learning rate determines how much the weights (weight parameter α) are learned in one training session. i This value determines whether to modify it.
[0113]
number
[0114] (2) Normalization: Normalize according to the following formula 6.
[0115]
number
[0116] The above steps (1) and (2) are repeated M times. Here, M is called the number of epochs. This results in each α being an appropriate value that reflects the tagged data (equivalent to training data). i You can obtain this.
[0117] Figure 12 shows an example of the weight values calculated for each pipeline (first generation unit 521 to nth generation unit 524 → feature extraction unit 530). The recording control unit 54 (see Figure 9) associates these weight values with the pipeline and stores them in the weight DB 600 (see Figure 10).
[0118] In the above, the loss function L defined in Equation 3 above is the tag value vector t(t1, t2, ..., t) of the tagged data (equivalent to training data). 15 ) and the expected value y(vector y(y1,y2,...,y 15 An example is shown where the absolute value of the difference between )) is squared, but the definition of the error function is not limited to this and other definitions may be used. For example, it may be defined as binary cross-entropy, mean squared error, mean absolute error, the square root of the mean squared error (RMSE (root-mean-square error) or RMSD (root-mean-square deviation)), mean squared logarithmic error (MSLE (Mean Squared Logarithmic Error)), Huber loss, Poisson loss, hinge loss, Kullback-Leibler divergence, etc.
[0119] Furthermore, the weights α of the weighted average are as follows: i An example was shown in which the initial value of is determined based on Equation 2 described above. However, the weight of the weighted average α i Other values may be set as the initial value. For example, a non-uniform weight α obtained by matching the tag information FQ1 to FQ4 generated by the feature extraction unit 530 with the tag information TD3 that constitutes the tagged data TD1. iYou may use this as the initial value. For example, the initial value dependency decreases with each iteration, but if you set the initial value close to the goal value, the closer you are to the goal value, the faster you can reach the goal. Therefore, by setting the initial value close to the goal value, it becomes possible to execute calculations quickly. This makes it possible to reduce the amount of calculation. In Figure 13, the weights of the weighted average α i This example shows how to set the initial value of a parameter to a value close to the goal value.
[0120] Figure 13 shows the weights α of the weighted average in the weight calculation process performed by the information processing device 500. i This diagram schematically illustrates an example of calculations when setting the initial value.
[0121] As shown in Figure 13(A), the weight calculation unit 540 calculates the difference value of features for each tag between the tag information FQ1, which is based on the text information TX1 generated by the first generation unit 521, and the tag information TD3. For example, the feature value "0.4" for the tag "casual" in tag information FQ1 is compared with the feature value "0.3" for the item "casual" in tag information TD3, and the difference value "0.1" is calculated. Difference values are calculated similarly for other tags. Then, the weight calculation unit 540 calculates a total value by summing the absolute values of the difference values of features calculated for each tag in tag information FQ1 and tag information TD3. Figure 13(A) shows an example where the total value calculated is "0.7".
[0122] Furthermore, as shown in Figure 13(B), the weight calculation unit 540 calculates the difference in feature quantities for each tag between the tag information FQ2, based on the text information TX2 generated by the second generation unit 522, and the tag information TD3. Also, similar to the example shown in Figure 13(A), the weight calculation unit 540 calculates the difference in feature quantities for each tag in the tag information FQ2 and tag information TD3, and calculates a total value by summing the absolute values of these difference quantities. Figure 13(B) shows an example where "1.2" is calculated as the total value. Although not shown in the diagram, similarly, the difference in feature quantities between the tag information FQ3 and FQ4, based on the text information TX3 and TX4 generated by the other generation units (third generation unit 523, nth generation unit 524), and the tag information TD3 is calculated, and the sum of the absolute values of these difference quantities is calculated.
[0123] Figure 13(C) shows the sum of the absolute values of the difference values obtained by each of the calculation processes described above. In this case, the weight calculation unit 540 calculates the weight α based on the magnitude of each sum. i It is possible to determine the initial values of the weights α as the total value increases. i It is possible to set a small initial value for each weight α. However, as with Equation 2 above, i Each weight α is such that the sum is 1. i Set the initial value of the weight α. i The calculation process after setting the initial values is the same as described above, so the explanation is omitted here. Furthermore, if there are multiple tagged data (equivalent to training data), it is possible to use the values obtained by applying a predetermined calculation (e.g., averaging) to each difference value calculated for each of those tagged data.
[0124] [Example of information processing device operation] Figure 14 is a flowchart illustrating an example of weight calculation processing in the information processing device 500. This weight calculation processing is performed based on a program stored in the storage unit 55. This weight calculation processing is also performed when tagged data TD1 is input to the acquisition unit 510 by user operation. This weight calculation processing will be explained with reference to Figures 9 to 13 as appropriate.
[0125] In step S701, the acquisition unit 510 receives the tagged data TD1 input by user U1, outputs the image information TD2 constituting the received tagged data TD1 to the text information generation unit 520, and outputs the tag information TD3 constituting the tagged data TD1 to the weight calculation unit 540.
[0126] In step S702, each generation unit (first generation unit 521 to the nth generation unit 524) constituting the text information generation unit 520 generates text information TX1 to TX4 (see Figure 11) based on the image information TD2 output from the acquisition unit 510.
[0127] In step S703, the feature extraction unit 530 generates tag information FQ1 to FQ4 (see Figure 11) based on the text information TX1 to TX4 generated by the text information generation unit 520.
[0128] In step S704, the weight calculation unit 540 calculates the weight values for each generation unit (first generation unit 521 to the nth generation unit 524) (weight values for each pipeline) based on the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the tag information TD3 that constitutes the tagged data TD1. This calculation process is the same as the calculation process described above, so its explanation is omitted here.
[0129] In step S705, the recording control unit 54 associates the weights calculated in step S704 for each tag information FQ1 to FQ4 (each pipeline) with each generation unit (first generation unit 521 to nth generation unit 524) that constitutes the text information generation unit 520 and stores them in the weight DB 600. For example, as shown in Figure 10, each piece of information is stored in the weight DB 600.
[0130] [Example of generating tag information using weights] Next, we will show an example of generating tag information using the weights described above. Figures 15 and 16 show an example of generating tag information by performing a weighted average calculation.
[0131] [Example of an information processing device configuration] Figure 15 is a block diagram showing an example of the functional configuration of the information processing device 560. Note that the information processing device 560 is a modified version of the information processing device 500 shown in Figure 9. Except for the addition of an image acquisition unit 51 and a weighted average calculation unit 550 (replacing the acquisition unit 510 and weight calculation unit 540), and the addition of an image database 200, it is identical to the information processing device 500. Therefore, parts common to both the information processing device 500 and the information processing device 560 are given the same reference numerals, and some of their descriptions are omitted.
[0132] The information processing device 560 comprises an image acquisition unit 51, a text information generation unit 520, a feature extraction unit 530, a weighted average calculation unit 550, a recording control unit 54, and a storage unit 55. The image acquisition unit 51, text information generation unit 520, feature extraction unit 530, weighted average calculation unit 550, and recording control unit 54 are each implemented by, for example, one or more processing circuits such as CPUs and GPUs. Also, although the information processing device 10 and the information processing device 560 are shown separately, as in the example shown in Figure 2, they may be configured as a single integrated device.
[0133] The weighted average calculation unit 550 performs a weighted average calculation based on the feature quantities for each tag related to the multiple text information output from the feature extraction unit 530 and the weight values stored in the weight DB 600. The weighted average calculation unit 550 then outputs the feature quantities for each tag related to the image (image acquired by the image acquisition unit 51) calculated by the weighted average calculation to the recording control unit 54. Specifically, the tag information FQ1 to FQ4 generated by the feature extraction unit 530 and the weight α stored in the weight DB 600 are output. i A weighted average calculation is performed using α. For example, for the tag "casual", each feature of the tag information FQ1 to FQ4 is weighted α i The sum of the values obtained by multiplying by α is used as the feature of the tag "casual". Similarly, for other tags, the feature of tag information FQ1~FQ4 and the weight α are used. i The sum of the values obtained by multiplying these two factors by themselves is used as the feature quantity for that tag.
[0134] [Example of information processing device operation] Figure 16 is a flowchart illustrating an example of feature extraction processing in the information processing device 560. This feature extraction processing is executed based on a program stored in the memory unit 55. This example shows a feature extraction process using a weighted average. This feature extraction processing is executed when an image is acquired by the image acquisition unit 51. The weight calculation process will be explained with reference to Figures 1 to 15 as appropriate.
[0135] In step S711, the image acquisition unit 51 acquires the image input by user U1 and outputs image information related to the acquired image to the text information generation unit 520.
[0136] In step S712, each generation unit (first generation unit 521 to the nth generation unit 524) constituting the text information generation unit 520 generates multiple text information based on the image information output from the image acquisition unit 51, and outputs the generated multiple text information to the feature extraction unit 530. The method for generating these multiple text information is the same as in the example shown in Figure 11.
[0137] In step S713, the feature extraction unit 530 extracts feature quantities for each of the multiple tags stored in the tag DB 100, based on each of the multiple text information generated by the text information generation unit 520. The feature extraction unit 530 then outputs the feature quantities for each of the multiple tags extracted for each text information to the weighted average calculation unit 550. The method for extracting feature quantities for each tag is the same as in the example shown in Figure 11.
[0138] In step S714, the weighted average calculation unit 550 performs a weighted average calculation on the feature quantities for each of the multiple tags extracted for each piece of text information by the feature extraction unit 530 to obtain a set of tag information (feature quantities for each of the multiple tags). The weighted average calculation unit 550 then outputs the obtained feature quantities for each of the multiple tags to the recording control unit 54.
[0139] In step S715, the recording control unit 54 stores the feature quantities for each of the multiple tags output from the weighted average calculation unit 550 and the image information output from the image acquisition unit 51 in the image DB 200. For example, as shown in Figure 4, the feature quantities for each of the multiple tags and the image information are associated and stored in the image DB 200.
[0140] [Example of the effects of the second embodiment] As mentioned above, using a single text information generation model (e.g., img2txt) results in a high degree of dependence on the features of that model. However, using multiple text information generation models (e.g., img2txt) can mitigate this dependence. Furthermore, using multiple models makes it possible to stabilize the behavior. Additionally, for example, generating a trained model to extract image features from a specific perspective based on user preferences requires a large amount of training data where features have been associated through human work. In contrast, in the second embodiment, it is possible to calculate weights using only one or a small amount of tagged data (equivalent to training data). In other words, in the second embodiment, it is possible to extract appropriate features even with a small amount of data.
[0141] [Example of modifying text information using weights] The above example demonstrates how to perform a weighted average calculation on the feature quantities for each tag based on multiple text information generated by multiple first generation units 521 to the nth generation unit 524, using weights calculated using one or a small amount of tagged data (corresponding to training data). However, these weights may be used in other ways to calculate the feature quantities for each tag. For example, weights may be used for multiple text information generated by each of the multiple first generation units 521 to the nth generation unit 524. For example, it is assumed that text information generated by a generation unit with a large weight has relatively high accuracy, and text information generated by a generation unit with a small weight has relatively low accuracy. Therefore, for example, text information generated by a predetermined number (e.g., 1 to 3) of generation units with large weights may be used as reference text information, and instruction information to modify one or more reference text information using text information generated by other generation units, and to extract the feature quantities for each tag from the modified one or more text information, may be output to the feature extraction unit 530, and the feature quantities for each tag may be extracted from the modified one or more text information. This makes it possible to generate corrected text information with even higher accuracy by modifying the base text information, which is estimated to have high accuracy, with other text information. Based on this corrected text information, it is possible to extract tag-specific features with even higher accuracy.
[0142] In other words, for example, for each of the text information TX1 to TX4 (see Figure 11) generated by the generation units (first generation unit 521 to the nth generation unit 524) that constitute the text information generation unit 520, new text information may be generated considering the weights, and feature quantities for each tag may be extracted based on that new text information.
[0143] For example, the importance of each sentence is set according to its weight for the first sentence contained in text information TX1, the second sentence contained in text information TX2, the third sentence contained in text information TX3, and the fourth sentence contained in text information TX4. For example, consider the case where the first generation unit 521 has the highest weight, the second generation unit 522 has the second highest weight, the third generation unit 523 has the third highest weight, and the Nth generation unit 524 has the lowest weight. In this case, the feature extraction unit 530 can generate a new sentence using the first sentence as a reference and referring to the second to fourth sentences. Therefore, for example, it is possible to output text information TX1 to TX4 to the feature extraction unit 530, and at the same time output instruction information to the feature extraction unit 530 instructing it to generate a new sentence using the first sentence as a reference and referring to the second to fourth sentences, and to extract features from that new sentence.
[0144] As a result, the weighted average calculation unit 550 can be omitted, and tag information can be generated using two or more text information generation units 520 and one feature extraction unit 530, taking weights into consideration. Furthermore, if there are multiple corrected text information items, it is possible to generate feature quantities for each set of tags by performing a weighted average calculation process using the weights.
[0145] [Example of setting tags based on text information] In the first embodiment, an example was shown of setting multiple tags based on text information generated from an image. Here, an example is shown of setting multiple tags based on at least one of the multiple text information generated by the multiple first generation units 521 to the nth generation unit 524. As mentioned above, it is assumed that text information generated by generation units with large weights has relatively high accuracy, and text information generated by generation units with small weights has relatively low accuracy. Therefore, for example, tags may be set based on the strings, words, etc. contained in one or more text information generated by a predetermined number (e.g., 1 to 3) of generation units with large weights. Note that the method for setting tags using text information is the same as the example shown in the first embodiment (an example of setting multiple tags based on text information generated from an image), so the explanation is omitted here.
[0146] [Differentiation] In the second embodiment, an example was shown in which tag information is generated using two or more types of text information generation units 520 and one type of feature extraction unit 530. However, tag information may also be generated using both two or more types of text information generation units and two or more types of feature extraction units. In this case, it is possible to calculate weights for each combination of text information generation units and feature extraction units and generate tag information using these weights.
[0147] Furthermore, the above example shows how to generate tag information while considering weights using two or more text information generation units 520, one feature extraction unit 53, and a weighted average calculation unit 550. However, the weighted average calculation unit 550 may be omitted, and tag information may be generated using two or more text information generation units 520 and one feature extraction unit 53. For example, text information obtained by combining each sentence contained in the text information TX1 to TX4 (see Figure 11) generated by the generation units (first generation unit 521 to nth generation unit 524) constituting the text information generation unit 520 may be output to the feature extraction unit 530.
[0148] For example, it is possible to output a combined sentence to the feature extraction unit 530 by simply arranging the first sentence from the first to the fourth sentence, using the first sentence contained in text information TX1, the second sentence contained in text information TX2, the third sentence contained in text information TX3, and the fourth sentence contained in text information TX4. In this case, the feature extraction unit 530 can extract features based on the combined sentence which is an arrangement of the first to the fourth sentences. Alternatively, for example, the combined sentence may be output to the feature extraction unit 530, and instruction information to perform a predetermined process on the combined sentence may also be output to the feature extraction unit 530. For example, it is possible to extract common parts (e.g., strings of words, sentences, etc.) from the combined sentence (first combined sentence), generate a new combined sentence (second combined sentence) by combining the extracted parts with non-common parts (e.g., strings of words, sentences, etc.), and output instruction information to instruct the extraction of features from the second combined sentence. Furthermore, it is possible to output instruction information that, for example, generates a new sentence based on the combined sentence (first combined sentence) according to some criteria, and then instructs the system to extract features from that new sentence (fifth sentence).
[0149] [Differentiation] Figures 1 and 2 show an example where the information processing device 10 that generates tags and the information processing device 50 that generates tag information using those tags are different devices. However, as mentioned above, the device that generates tags and the device that generates tag information using those tags may be the same device. Furthermore, although the above examples show the generated tag information being stored in the image DB 200 in association with an image, the generated tag information may also be displayed together with the image, or transmitted to other devices in association with the image. Therefore, Figures 17 and 18 show an example where the processing unit that generates tags and the processing unit that generates tag information using those tags are the same device. Also, Figures 17 and 18 show an example where the generated tag information is displayed together with the image.
[0150] [Example of an information processing device configuration] Figure 17 is a block diagram showing an example of the functional configuration of the information processing device 800.
[0151] The information processing device 800 is an example of an information processing device comprising the configuration of the information processing device 10 shown in Figure 1 and the configuration of the information processing device 50 shown in Figure 2. Parts common to both the information processing device 10 shown in Figure 1 and the information processing device 50 shown in Figure 2 are denoted by the same reference numerals as those for the information processing devices 10 and 50, and some of their descriptions are omitted.
[0152] The information processing device 800 comprises a DB control unit 810 and an output unit 820. The DB control unit 810 performs recording control and reading control to each DB of the storage unit 55. For example, the DB control unit 810 stores the tag list output from the tag setting unit 12 in the tag DB 100. The DB control unit 810 also associates the feature quantities for each tag output from the feature extraction unit 53 with the images output from the image acquisition unit 51 and stores them in the image DB 200. Furthermore, the DB control unit 810 associates the image information and tag information stored in the image DB 200 and outputs them from the output unit 820. An example of this output will be explained in detail with reference to Figure 18.
[0153] The output unit 820 is an output unit that outputs various types of information (for example, a display unit 821 (see Figure 18) and an audio output unit (not shown)). Specifically, the output unit 820 displays various types of information on the display unit 821 and outputs audio information from the audio output unit based on the information output from the DB control unit 810. An image display device such as a display screen capable of outputting images and audio can be used as the output unit 820. Alternatively, for example, an output device capable of outputting at least one of images and audio may be used. In Figure 17, an example of the output unit 820 is shown in which the display unit 821 and the audio output unit are provided on the information processing device 800, but a separate output device different from the information processing device 800 may be used as the output unit 820.
[0154] [Example of tag information display] Figure 18 shows an example of how tag information generated by the information processing device 800 is displayed on the display unit 821. As shown in Figure 18, it is possible to display the image MG1 acquired by the image acquisition unit 51 and the feature quantities (tag information TG10) extracted by the feature extraction unit 53 in association with each other on the display unit 821. This makes it possible for user U1 to easily check the tag information TG10 generated for a desired image MG1.
[0155] [Example of a communication system configuration] The above example demonstrates how to extract feature quantities for each tag using equipment accessible to user U1. However, various input operations may be performed on the first device, and various processing (e.g., feature extraction processing for each tag) corresponding to those input operations may be performed on the second device. Furthermore, the various processing according to this embodiment may be performed using three or more devices. Therefore, Figures 19 and 20 show an example of a communication system TS1 that can exchange various information using multiple devices connected via a network NW1.
[0156] Figure 19 is a block diagram showing an example of the functional configuration of the communication system TS1.
[0157] The communication system TS1 consists of a network NW1, information processing devices 900 and 920, electronic equipment 930, etc. For example, each of the network NW1, information processing devices 900 and 920, and electronic equipment 930, etc., is connected via network NW1. Communication between these devices is carried out using either wired communication or wireless communication. In addition, communication between these devices may be carried out directly, rather than via network NW1.
[0158] Network NW1 is a network such as a public telephone network or the Internet. Furthermore, each device constituting the communication system TS1 is connected to Network NW1 by either a wireless communication method, a wired communication method, or both.
[0159] The information processing device 900 corresponds to the information processing device 800 shown in Figure 17. Furthermore, each part of the information processing device 900 corresponds to the same part shown in Figure 17. The information processing device 900 can, for example, function as a server capable of providing various types of information.
[0160] The communication unit 910 exchanges various types of information with other devices using at least one of wired communication or wireless communication. For example, the communication unit 910 performs receiving processing to receive various types of information transmitted from the information processing device 920 and the electronic device 930, and transmitting processing to send various types of information to the information processing device 920 and the electronic device 930.
[0161] The information processing device 920 and the electronic device 930 are fixed or portable information processing devices owned by user U1, such as smartphones, tablet terminals, smartwatches, and personal computers. Furthermore, the information processing device 920 and the electronic device 930 are capable of wired or wireless communication with the information processing device 900. The information processing device 920 and the electronic device 930 also perform transmission processing to send various information to the information processing device 900, and reception processing to receive various information from the information processing device 900. For example, based on the information received from the information processing device 900, the information processing device 920 and the electronic device 930 can display various images on the display units 921 and 931, or output various sounds.
[0162] Figure 20 is a diagram showing an example of the transition of the display screen displayed on the display unit 931 of the electronic device 930.
[0163] Figure 20(A) shows an example of the display screen 932 that appears when setting tags and images to be tagged. The display screen 932 shows an input field display area 933, a confirmation button 934, and an image list display area MG10. The information entered in the input field display area 933 is the same as the information entered in the text input field 17 shown in Figure 5(A). The images displayed in the image list display area MG10 and selected are the same as the image group 40 shown in Figure 2, etc. For example, if user U1 enters the string "clothes" into the input field display area 933, and then selects an image desired by user U1 from the images displayed in the image list display area MG10, and then selects the confirmation button 934, then the text information related to the string "clothes" and the image information related to the selected image are sent to the information processing device 900. When the information processing device 900 receives this information, it generates tag information for the selected image based on this information. The method for generating this tag information is the same as the generation method described above. The information processing device 900 then performs display control to display the generated tag information and the image information related to the selected image on the electronic device 930. An example of this display is shown in Figure 20(B).
[0164] Figure 20(B) shows an example of the display screen 935, which shows the image MG1 and tag information TG10 in association. Note that the image MG1 and tag information TG10 are the same as in the example shown in Figure 18. In this way, the user U1 can easily obtain the image's tag information using a portable device or similar.
[0165] Note that while Figures 17 and 19 show an example configuration combining the configuration shown in Figure 1 and the configuration shown in Figure 2, an example configuration combining the configuration shown in Figure 9 and the configuration shown in Figure 15 may also be used.
[0166] [Examples of application to videos] The above example illustrates the extraction of feature quantities from multiple tags for a still image. However, this embodiment can also be applied to videos composed of multiple images (frames). For example, by extracting feature quantities from multiple tags for images at predetermined intervals (e.g., images at intervals of several seconds to tens of seconds) among the multiple images (frames) that make up a video, and performing a predetermined calculation (e.g., calculating the average value for each tag) on the extracted feature quantities of multiple tags, it is possible to extract feature quantities from multiple tags related to the entire video. In this case, for example, an image representing the video (e.g., a representative image) and the extracted tag information (feature quantities from multiple tags) may be associated and stored in a video database, or they (image representing the video, extracted tag information) may be associated and displayed on a display unit, or they may be associated and transmitted to other devices. This makes it possible to easily grasp the tag information related to a single video. Furthermore, for example, in a video containing multiple scenes (e.g., a video in which the shooting scenes are changed sequentially), the predetermined calculation process described above may be performed for each scene to extract feature quantities from multiple tags related to each scene. In this case, for example, images representing each scene (e.g., representative images for each scene) and tag information extracted for each scene (feature quantities of multiple tags) may be associated and stored in a video database, or they may be associated and displayed on a display unit, or they may be associated and transmitted to other devices. This makes it possible to easily grasp the tag information for each scene related to a single video.
[0167] [Third Embodiment] [Example of searching for images desired by the user] The above example demonstrated a tagging process that extracts multiple tag features from an image and tags it accordingly. Images tagged in this way can be used to search for images that the user likes or desires. Therefore, the following section will explain an example of a search process for images that the user likes or desires.
[0168] [Example of an information processing system configuration] Figure 21 is a block diagram showing an example configuration of the information processing system IS1.
[0169] The information processing system IS1 consists of an information processing device 1000 and an electronic device 1100. For example, the information processing device 1000 and the electronic device 1100 are connected via a predetermined network NW1. Communication between these devices can be conducted using either wired communication or wireless communication. In addition, communication between these devices may be conducted directly, rather than via the network NW1.
[0170] [Example Server Configuration] The information processing device 1000 comprises a communication unit 1010, a control unit 1020, and a storage unit 1030. The information processing device 1000 can be implemented using information processing devices and electronic devices such as servers, personal computers, smartphones, and tablet terminals. The control unit 1020 can be implemented using, for example, one or more processing circuits such as CPUs and GPUs.
[0171] The communication unit 1010 exchanges various types of information with other devices using at least one of wired communication or wireless communication. For example, the communication unit 1010 performs receiving processing to receive various types of information transmitted from the electronic device 1100, and transmitting processing to send various types of information to the electronic device 1100.
[0172] The control unit 1020 includes an information acquisition unit 1021, a tag generation unit 1022, a vector generation unit 1023, an extraction unit 1024, and an information provision unit 1025.
[0173] The information acquisition unit 1021 acquires various types of information transmitted from the electronic device 1100 and supplies the acquired information to each unit.
[0174] The tag generation unit 1022 generates tags for information (e.g., image information, text information) output from the information acquisition unit 1021, and outputs information about the generated tags (tag information) to the vector generation unit 1023. For example, if the image information output from the information acquisition unit 1021 is an image stored in the image information DB 1040, the tag generation unit 1022 acquires the feature quantity 1044 (see Figure 23) associated with that image. The identity of the image can be determined based on the identification information associated with the image.
[0175] Furthermore, if the information output from the information acquisition unit 1021 is other image information (images not stored in the image information DB 1040) or text information, the tag generation unit 1022 can generate tags based on tag information stored in the tag information DB 1060. Also, if the information output from the information acquisition unit 1021 is other image information or text information, the tag generation unit 1022 can generate tags using an AI (Artificial Intelligent) model. The method of tag generation will be explained in detail with reference to Figures 25 to 29, etc.
[0176] The vector generation unit 1023 generates user vectors according to the user's preferences for the electronic device 1100 based on the tag information (vector components) output from the tag generation unit 1022, and outputs the generated user vectors to the extraction unit 1024. The method for generating user vectors will be explained in detail with reference to Figures 25 to 29, etc.
[0177] The extraction unit 1024 extracts desired information from each DB stored in the storage unit 1030 based on the information output from each unit, and outputs the extracted information to the information provision unit 1025. For example, when a user vector is output from the vector generation unit 1023, the extraction unit 1024 compares that user vector with the provider vector (tag 1053, feature quantity 1054 (see Figure 24)) stored in the provider information DB 1050, and extracts provider information according to the user's preferences of the electronic device 1100 based on the similarity of each vector. Also, for example, when the electronic device 1100 requests the provision of a new image, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with the feature quantity 1044 (vector component) stored in the image information DB 1040, and extracts an image according to the user's preferences of the electronic device 1100 based on the similarity of each vector. The methods for generating each of these user vectors will be explained in detail with reference to Figures 25 to 29, etc.
[0178] The information provision unit 1025 performs transmission control to send the information output from each unit to the electronic device 1100.
[0179] The memory unit 1030 is a storage medium for storing various types of information. For example, the memory unit 1030 stores various types of information necessary for the control unit 1020 to perform various processing (e.g., control program, image information 1040 (see Figure 23), provider information DB 1050 (see Figure 24), tag information DB 1060, AI model). The memory unit 1030 also stores various types of information acquired via the communication unit 1010. The memory unit 1030 can be, for example, ROM, RAM, SRAM, HDD, SSD, or a combination thereof.
[0180] Image information DB1040 is a database for storing images provided by providers and associating them with the feature quantities of multiple tags generated based on these images. Here, a provider refers to, for example, an individual or company that provides various goods, services, etc., using a website (e.g., an e-commerce site). Image information DB1040 will be explained in detail with reference to Figure 23.
[0181] The provider information DB 1050 is a database for storing and associating various information about providers who provide various goods, services, etc., to users of the electronic device 1100. The provider information DB 1050 will be explained in detail with reference to Figure 24.
[0182] The tag information DB 1060 is a database for storing information used when tag information is generated by the tag generation unit 1022. For example, when tag information is generated using a dictionary database that associates strings, tags corresponding to those strings, and various other pieces of information including features corresponding to those tags, information about those various pieces of information is stored in the DB 1060. Also, for example, when tag information is generated using an AI model, information about the AI model is stored in the DB 1060.
[0183] [Example of electronic device configuration] The electronic device 1100 comprises a communication unit 1110, a location information acquisition unit 1120, an image acquisition unit 1130, a sound acquisition unit 1140, a control unit 1150, a UI (User Interface) unit 1160, and a storage unit 1170. The electronic device 1100 can be implemented by devices such as smartphones, tablet terminals, personal computers, or other information processing devices and electronic equipment. Note that the components of the electronic device 1100 are just examples, and some of them may be omitted, or other components may be added. For example, if the electronic device 1100 is a personal computer, the location information acquisition unit 1120 and the image acquisition unit 1130 can be omitted or replaced with external devices.
[0184] The communication unit 1110, based on the control of the control unit 1150, exchanges various types of information with other devices using wired or wireless communication.
[0185] The location information acquisition unit 1120 acquires location information relating to the location of the electronic device 1100 and outputs the acquired location information to the control unit 1150. The location information acquisition unit 1120 can be implemented, for example, by a GPS receiver that receives GPS (Global Positioning System) signals and calculates location information based on those GPS signals. The calculated location information includes data related to the location, such as latitude, longitude, and altitude at the time of receiving the GPS signal. Alternatively, a location information acquisition device that acquires location information using other methods may be used. For example, a location information acquisition device that derives location information using access point information from surrounding wireless LANs and acquires this location information may be used. Alternatively, a location information acquisition device that derives location information using location information from base stations used for calls and communications and acquires this location information may be used. Alternatively, a location information acquisition device that derives location information using location estimation technology based on navigation functions and acquires this location information may be used. For example, the location of a device can be estimated based on sensor information from various sensors (e.g., accelerometers, gyroscopes) and map information.
[0186] The image acquisition unit 1130 captures an image of a subject and generates an image (image data) based on the control of the control unit 1150, and outputs the generated image to the control unit 1150. The image acquisition unit 1130 is composed of, for example, an image sensor that receives light from a subject focused by a lens (not shown), and an image processing unit that performs predetermined image processing on the image data generated by the image sensor. For example, a CCD (Charge Coupled Device) type or a CMOS (Complementary Metal Oxide Semiconductor) type image sensor can be used as the image sensor.
[0187] The sound acquisition unit 1140 acquires ambient sounds around the electronic device 1100 based on the control of the control unit 1150, and outputs sound information related to the acquired sounds to the control unit 1150. For example, one or more microphones or sound acquisition sensors can be used as the sound acquisition unit 1140.
[0188] The control unit 1150 controls each part based on various programs stored in the memory unit 1170. The control unit 1150 is implemented by a processing unit such as a CPU or GPU. The processes performed by the control unit 1150 will be explained in detail with reference to Figures 26 to 29, etc.
[0189] The memory unit 1170 is a storage medium for storing various types of information. For example, the memory unit 1170 stores various types of information necessary for the control unit 1150 to perform various processes (e.g., control programs, image information DB 1080 (see Figure 33), search applications, device identification information). The memory unit 1170 also stores various types of information acquired via the communication unit 1110. The memory unit 1170 can be, for example, a ROM, RAM, SRAM, HDD, SSD, or a combination thereof. The memory unit 1170 may be a removable memory unit or a memory unit built into the electronic device 1100. The electronic device 1100 may also be provided with both a removable memory unit and a built-in memory unit.
[0190] The UI unit 1160 includes a reception unit 1161 and an output unit 1162.
[0191] The reception unit 1161 receives various operations from the user and outputs the received operation details to the control unit 1150. The reception unit 1161 and output unit 1162 may be configured as a touch panel that allows the user to input operations by touching or bringing their finger close to the display surface, or they may be configured as a separate user interface. When configured as a separate user interface, various operating components such as switches, buttons, and keyboards can be used as the reception unit 1161. Alternatively, an image corresponding to the operating unit may be projected onto a wall or the like, and the user may use that image (for example, by pointing at it) to perform operations. Furthermore, a sound acquisition unit 1140 or the like that receives various operations based on the user's voice may be used as the reception unit 1161.
[0192] The output unit 1162 outputs various information based on the control of the control unit 1150. For example, the output unit 1162 can be composed of a display unit and a sound output unit. The display unit displays various images based on instructions from the control unit 1150. As the display unit, for example, a display panel such as an organic EL (Electro Luminescence) panel or an LCD (Liquid Crystal Display) panel can be used. The sound output unit outputs various sounds based on instructions from the control unit 1150. As the sound output unit, for example, one or more speakers can be used. Note that the display unit, sound output unit, sound acquisition unit 1140, and reception unit 1161 are examples of a user interface, and some of these may be omitted, or other user interfaces may be used.
[0193] [Tagging examples] This section describes the tagging process applied to image information stored in the image information database 1040. For example, image information uploaded to the information processing device 1000 by a provider using the service in this embodiment is stored in the image information database 1040. That is, image information uploaded to the information processing device 1000 by a provider is registered as a provided image. When an image is registered in this way, a tagging process is performed on that image. The registered image and the tag generated for it are then associated and stored in the image information database 1040.
[0194] Here, as mentioned above, a tag is a criterion for classifying images based on the properties of each image, and also refers to the criteria used in that classification. In other words, tagging an image can be called assigning meaning to an image. Furthermore, a tag refers to the criteria used when extracting the features (feature quantities) of an image. Tags can also be called information for identifying an image. Note that the labels shown below also refer to information that represents the features, attributes, etc. of an image. Thus, since tags and labels have almost the same or similar meanings, in the following, tags and labels may be used interchangeably.
[0195] Furthermore, in this embodiment, as described above, information relating tags to their corresponding features (feature quantities) will be referred to as tag information. However, tag information may also be referred to as a tag dictionary, metadata, supplementary information, additional information, etc. In addition, in this embodiment, an example in which numerical feature quantities are used as features corresponding to tags will be described.
[0196] For example, tags such as simple, modern, retro, classic, casual, fancy, natural, Art Deco, bohemian, floral, animal motif, ethnic, and Nordic can be used as tags attached to an image of a mug (i.e., tags that represent the design style of the mug).
[0197] Each of these tags can be represented, for example, as presence or absence (e.g., 0 or 1), or as a multi-level representation (e.g., 5 levels from 0 to 4, or a range from 0 to 1). Figure 23 shows an example of each tag being represented as a multi-level representation (range from 0 to 1).
[0198] [Example of tagging performed manually] When an image is uploaded to the information processing device 1000 by a provider, it is possible to perform a tagging process in which a human is assigned tags to that image. For example, a human can judge the elements of the objects contained in the image (e.g., the style and taste of the objects contained in the image), and based on the result of that judgment, the aforementioned tag information can be assigned to the image. However, when tags are assigned in this way by human work, personnel with specialized knowledge are required at the time of registration, and it is often difficult to obtain such personnel. Therefore, the following shows an example of performing the tagging process using artificial intelligence (AI), etc. Note that in this embodiment, an example of assigning tags using AI will be mainly described.
[0199] [Example of tagging using AI] [Example of images and training labels used for training] Figure 22(A) shows an example of image IMG1 used for training. Here, we will use an image of a mug as an example of image IMG1. In other words, when performing a tagging process that assigns tags to images using AI, various images such as image IMG1 can be used as training data. Note that image IMG1 is a simplified image for the sake of clarity.
[0200] Alternatively, a tag pattern may be prepared for each image category (e.g., mug, sofa furniture, house), and an AI model may be generated for each image category, or a common tag pattern may be prepared regardless of the image category, and an AI model may be generated. Figure 22(B) shows an example of tags to be assigned as training labels to the image IMG1 used for training.
[0201] As shown in Figure 22(B), when training image IMG1, predetermined tags are assigned to image IMG1 as training labels. Then, training is performed using these assigned tags.
[0202] Figure 22(B) also shows an example of using tags that quantify the degree to which something fits on a scale of 0 to 1. In this case, a tag value of 0 means the degree of the relevant content is the lowest, and a tag value of 1 means the degree of the relevant content is the highest. For example, in Figure 22(B), the value (feature quantity F1) for tag TI1 "Art Deco" is set to "0.9". Therefore, this means that image IMG1 (mug) has a very high degree of "Art Deco". In other words, the mug contained in image IMG1 gives a strong impression of an Art Deco style.
[0203] Similarly, other images besides image IMG1 used for training are also assigned predetermined tags as training labels and used for training.
[0204] [Example of generating an AI model for training] An AI model (for example, a machine learning model generated through machine learning) is generated using the aforementioned training data (e.g., image IMG1) and training labels (e.g., tag TI1, feature F1). For example, it is possible to pre-train the AI model by loading the training data (e.g., image IMG1) and training labels (e.g., tag TI1, feature F1), and then use this AI model for tagging. As the AI model, for example, CNN (Convolutional Neural Network), ViT (Vision Transformer), etc., can be used. These can be executed by installing a predetermined application on the information processing device 1000.
[0205] [Example output from a trained AI model] As described above, when preparing tag patterns for each image category (e.g., mug, sofa furniture, house) and generating an AI model for each image category, multiple AI models are prepared for each tag pattern. In this case, the image registered in the information processing device 1000 is input to the AI model corresponding to the image category. As a result, the AI model corresponding to the image category outputs the result with tags corresponding to that category. In this way, it is possible to output the input image with tag information related to the set tags using multiple AI models. Also, as described above, when preparing tag patterns and generating an AI model regardless of the image category, one AI model is prepared. In this case, the image registered in the information processing device 1000 is input to that AI model. As a result, that AI model outputs the result with tags. In this way, it is possible to output the input image with tag information related to the set tags using a single AI model.
[0206] [Example of image information database configuration] Figure 23 is a simplified diagram showing the contents of the image information stored in the image information DB 1040. In Figure 23, an example of storing image information for a mug is shown as an example of an image category. Note that in Figure 23, for the sake of ease of explanation, an image information DB is prepared for each image category (e.g., mug, furniture sofa, house), and image information is stored for each image category, but this is not the only example. For example, one or more common image information DBs may be prepared and various types of image information may be stored regardless of the image category.
[0207] Image information DB1040 is a database for storing images provided by providers and associating them with the feature quantities of multiple tags generated based on those images.
[0208] Specifically, image information 1042, tag 1043, feature quantity 1044, and provider identification information 1045 are associated with image identification information 1041 and stored in the image information DB 1040.
[0209] Image identification information 1041 is identification information used to identify each image prepared by the provider for the user. In Figure 23, for the sake of simplicity, an example is shown in which only numerical values are stored as identification information in image identification information 1041, but other types of identification information may also be used.
[0210] Image information 1042 is information about an image prepared by the provider for the user. Note that Figure 23 shows an example where only the image is stored in image information 1042 for ease of explanation, but various attribute information associated with the image may also be included.
[0211] Tag 1043 is information indicating the tags assigned to the image stored in image information 1042. For example, if the image stored in image information 1042 is an image of a mug, then as mentioned above, tags such as simple, modern, retro, classic, casual, fancy, natural, art deco, bohemian, floral, animal motif, ethnic, and Nordic will be stored.
[0212] Feature quantity 1044 is information indicating the features assigned to each tag for the image stored in image information 1042. Specifically, it stores the tag information assigned by the tag assignment process described above.
[0213] For example, when displaying images stored in the image information DB 1040 on the UI unit 1160 (see Figure 21) of the electronic device 1100, it is possible to use feature vector 1044 to narrow down the images to the user's preference. For example, if a user has a preference for fancy designs, it is possible to display search results on the UI unit 1160 that are narrowed down to images with a high value of "fancy" based on feature vector 1044 associated with the tag 1043 "fancy". In this case, the user can select their preferred image from among the images displayed on the UI unit 1160 (for example, images in which feature vector 1044 corresponding to tag 1043 "fancy" has a value of a predetermined value or higher). This makes it possible to generate a user vector (described later) that better reflects the user's preferences. Also, for example, if a user wants to be recommended only by providers that handle Art Deco style mugs, it is possible to display search results on the UI unit 1160 that are narrowed down to images with a high value of "Art Deco" based on feature vector 1044 associated with the tag 1043 "Art Deco". In this case, the user can select their preferred image from among the images displayed in the UI unit 1160 (images with a high "Art Deco" value). This makes it possible to more appropriately recommend providers that handle Art Deco style mugs. Note that the details of narrowing down the search for these images and provider information are omitted from the illustrations and other explanations.
[0214] The provider identification information 1045 is identification information for identifying the provider who provided the image stored in the image information 1042. In Figure 23, for the sake of clarity, an example is shown in which the provider's name is stored in the provider identification information 1045, but other information that can identify the provider (e.g., serial number, code) may also be used.
[0215] Note that the information shown in Figure 23 is just an example of the information stored in the image information DB 1040. Other information may be stored in association with it, and some of this information may be omitted as needed.
[0216] [Example of provider information database configuration] Figure 24 is a simplified diagram showing the contents of the provider information stored in the provider information DB 1050.
[0217] The provider information DB1050 is a database for storing various types of information about providers in an associated manner.
[0218] Specifically, provider identification information 1051, profile information 1052, tags 1053, and features 1054 are associated and stored in the provider information DB 1050.
[0219] Provider identification information 1051 is identification information for identifying the provider. Note that provider identification information 1051 corresponds to provider identification information 1045 shown in Figure 23. Also, similar to the example in Figure 23, Figure 24 shows an example where the provider's name is stored in provider identification information 1051 for ease of explanation, but other information that can identify the provider (e.g., serial number, code) may also be used.
[0220] Profile information 1052 is various information about the provider whose provider identification information is stored in provider identification information 1051. For example, the provider's address and provider introduction information are stored. This profile information can be edited by the provider as needed.
[0221] Tag 1053 is information (tag information) indicating a tag assigned to a provider whose provider identification information is stored in provider identification information 1051. This tag information stores information about all or part of the tags 1043 shown in Figure 23.
[0222] Feature 1054 is information indicating the features assigned to a provider whose provider identification information is stored in provider identification information 1051. For example, features related to goods or services provided by the provider (or features related to images containing those goods or services) can be used as these features. These features may be set as appropriate by the provider, or they may be stored based on the calculation results of each tag information assigned to multiple images uploaded to the information processing device 1000 by the provider. For example, the average value of each feature related to each tag assigned to multiple images uploaded to the information processing device 1000 can be stored in the corresponding feature. These features are used when selecting a provider to recommend to the user.
[0223] Note that the information shown in Figure 24 is just an example of the information stored in the provider information DB 1050. Other information may be stored in association with it, and some of this information may be omitted as needed.
[0224] [Example of selecting your preferred image] Figure 25 shows an example of the display of the search screen 1200 shown on the UI unit 1160 of the electronic device 1100. The search screen 1200 is displayed on the output unit 1162 based on the control of the control unit 1150.
[0225] The search screen 1200 displays a text input area 1201 and a select button 1202. Specifically, the information provision unit 1025 of the information processing device 1000 transmits search screen information to the electronic device 1100 based on a request from the electronic device 1100, in order to display the search screen 1200. When the electronic device 1100 receives this search screen information, the control unit 1150 displays the search screen 1200 on the UI unit 1160 based on the received search screen information. The text input area 1201 allows the user to input desired characters based on user operations. Figure 25 shows an example where the text information "I want a stylish mug" is entered into the text input area 1201. When text information is entered into the text input area 1201 and the select button 1202 is selected, the control unit 1150 of the electronic device 1100 transmits search information including that text information to the information processing device 1000. When the information processing device 1000 receives the search information, the extraction unit 1024 performs a selection process to select an image to be provided to the electronic device 1100 from the image information DB 1040 based on the text information contained in the received search information. For this selection process, known selection techniques for selecting images from text information can be used. Alternatively, a selection process described later (see Figures 36 to 38) may be used. Images from other devices (e.g., an image provision server) may also be searched and used. An example of an image selected in this selection process is shown in Figure 26. Note that in Figure 25, for the sake of clarity, an example is shown in which an image information DB 1040 prepared for each image category (e.g., mug, sofa furniture, house) is used, but the explanation is not limited to this. For example, the same applies when performing a selection process to select an image to be provided to the electronic device 1100 using one or more image information DBs prepared independently of image categories.
[0226] Figure 26 shows an example of the display of the selection screen 1210 shown on the UI unit 1160 of the electronic device 1100. The selection screen 1210 is displayed on the output unit 1162 based on the control of the control unit 1150.
[0227] On the selection screen 1210, an image display area 1211 for displaying a plurality of images 1213 to 1220 and a decision button 1221 are displayed. Specifically, the information providing unit 1025 of the information processing apparatus 1000 transmits selection screen information for displaying the image selected by the selection process of the extraction unit 1024 to the electronic device 1100. When the electronic device 1100 receives the selection screen information, the control unit 1150 causes the UI unit 1160 to display the selection screen 1210 based on the received selection screen information. Note that other images can be displayed in the image display area 1211 based on a user operation (e.g., a scroll operation).
[0228] Also, when a user operation is received by the reception unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on the user operation. For example, when a selection operation (e.g., a touch operation) for selecting a favorite image displayed in the image display area 1211 is performed by the user, the favorite image on which the selection operation has been performed is set to a selected state. For example, the selected state is made visible to the user by setting the favorite image on which the selection operation has been performed to a display mode different from other images (e.g., attaching selection marks 1222, 1223 (star areas), surrounding with a thick line, attaching a specific color). FIG. 26 shows an example in which images 1218 and 1220 are set to the selected state by attaching selection marks 1222 and 1223. Note that after the selection operation by the user, only the image set to the selected state by the user may be displayed. Note that it is also possible to cancel the selected state of the image by performing a selection operation (e.g., a click operation) on the selection marks 1222, 1223 (star areas).
[0229] Also, after one or more images are selected by the user, when the user executes a selection operation (e.g., a touch operation) for selecting the decision button 1221, the control unit 1150 transmits image selection information regarding the one or more images selected by the user to the information processing apparatus 1000. When the information processing apparatus 1000 receives the image selection information, the information acquisition unit 1021 outputs the received image selection information to the tag generation unit 1022 and the vector generation unit 1023.
[0230] [Example of generating user vector] The tag generation unit 1022 acquires tag information associated with one or more images corresponding to the image selection information transmitted from the electronic device 1100 from the image information DB 1040 (see FIG. 23). Specifically, the tag generation unit 1022 acquires one or more tag information (feature amounts) associated with the image identification information 1041 corresponding to one or more images corresponding to the image selection information from the tags 1043 and feature amounts 1044 of the image information DB 1040, and outputs the acquired tag information to the vector generation unit 1023.
[0231] Next, the vector generation unit 1023 generates a user vector regarding the user who transmitted the image selection information based on the tag information output from the tag generation unit 1022. When there is one image corresponding to the image selection information transmitted from the electronic device 1100, the tag information associated with that one image can be used as the user vector.
[0232] For example, the vector generation unit 1023 can generate a user vector by performing an addition process on each element constituting each vector corresponding to the tag information output from the tag generation unit 1022 and dividing the addition result by the number of vectors that are the addition targets. In this embodiment, for the sake of easy explanation, an example is shown in which a common tag (tag 1043 (see FIG. 23)) for each image is set regardless of the category of the image. Similarly, an example of setting a common tag (tag 1053 (see FIG. 24)) for each provider is shown.
[0233] For example, let's assume that the user selects image 3. The elements (components) that make up the tag information (tag vector) associated with this selected image 3 are [a1, ... , an1], [b1, ... , bn1], [c1, ... , cn1]. Here, n1 represents the number of each item (each element) of the tag. That is, it represents the number of tags stored in tag 1043 (see Figure 23). For example, if the tags used are simple, modern, retro, classic, casual, fancy, natural, art deco, bohemian, floral, animal motif, ethnic, and Nordic, then n1 = 13. Also, the component (numerical value) of each element is the value of feature 1044 (see Figure 23). For example, if the value of feature 1044 is in the range of 0 to 1, then the component (numerical value) of each element is a value in the range of 0 to 1.
[0234] In this case, the vector generation unit 1023 generates a user vector [(a1+b1+c1) / 3,...,(an1+bn1+cn1) / 3] by adding the corresponding components of the three multidimensional vectors ([a1,...,an1],[b1,...,bn1]) and dividing the result by 3. In this case, each element of the user vector corresponds to each element of the provider vector (each element of feature quantity 1054 (see Figure 24)).
[0235] Thus, it is possible to use the average value of multiple images selected by the user as the user vector. However, this embodiment is not limited to this, and it is possible to generate the user vector using other calculation methods. For example, it is possible to employ a calculation method in which the softmax function is used for each tag of each image and then the average is calculated. Alternatively, for example, it is possible to employ a calculation method in which the Max function is used to calculate the average. Furthermore, although this embodiment shows an example of generating a user vector using the elements that constitute each tag, it is not limited to this, and it is also possible to generate a user vector using some of the elements of each tag, or to generate a user vector using other tag information.
[0236] [Example of provider selection] This section shows an example of selecting an image provider based on an image selected by the user. The extraction unit 1024 extracts providers from among those stored in the provider information DB 1050 that correspond to the user's preferences, based on the user vector generated by the vector generation unit 1023. For example, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with the provider vector stored in feature quantity 1054 of the provider information DB 1050, and extracts a provider that corresponds to the user's preferences based on this comparison result.
[0237] Here, we show an example where a provider vector (feature 1054) is represented as [S1, ... , Sn1]. Here, n1 represents the number of items corresponding to each tag. Furthermore, each element's component (numerical value) corresponds to the value of feature 1054.
[0238] In this case, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with a certain provider vector (feature quantity 1054) [S1, ... , Sn1], and determines the similarity between the two based on the comparison result. For example, if the user vector is [(a1+b1+c1) / 3, ... , (an1+bn1+cn1) / 3], this user vector is compared with a certain provider vector [S1, ... , Sn1].
[0239] For example, the extraction unit 1024 calculates the cosine similarity between the user vector and each provider vector for each provider vector. A known method can be used to calculate this cosine similarity. Furthermore, it is preferable to normalize each vector when calculating the cosine similarity. The extraction unit 1024 then extracts the provider corresponding to the provider vector with the highest cosine similarity (i.e., the provider with the smallest cosine value with the user vector) as a provider according to the user's preference. A high cosine similarity means that the vectors are close in distance in the vector space. When providing multiple providers to the user, the extraction unit 1024 can extract a predetermined number of providers (for example, 2 to 10) with high calculated cosine similarity. The number of providers extracted in this way may be specified based on user operation using the electronic device 1100.
[0240] The provider extraction process may also be performed using other calculation methods. For example, the extraction unit 1024 calculates the difference value (for each corresponding component) between each component constituting the user vector and each component constituting each provider vector (the component corresponding to each component of the user vector) for each provider vector, and calculates the sum of these difference values for each provider vector. The extraction unit 1024 may then extract the provider corresponding to the provider vector with the smallest sum of these difference values as a provider that matches the user's preferences.
[0241] Thus, one provider may be selected that corresponds to the provider vector most similar to the user vector, or multiple providers may be selected in descending order of similarity that correspond to a predetermined number of provider vectors with a high similarity to the user vector. In other words, one provider may be selected that corresponds to the provider vector closest to the user vector, or multiple providers may be selected in descending order of proximity that correspond to provider vectors that are close in distance.
[0242] Furthermore, the extraction unit 1024 outputs information about the extracted provider (for example, profile information 1052, image information 1042) to the information provision unit 1025.
[0243] The information provision unit 1025 then transmits the selected provider information (for example, profile information 1052, image information 1042) regarding the extracted providers to the electronic device 1100.
[0244] [Example of displaying information about the selected provider] Figure 27 shows an example of the display of the provider information screen 1230 shown on the UI unit 1160 of the electronic device 1100. The provider information screen 1230 is displayed on the UI unit 1160 based on the control of the control unit 1150.
[0245] The provider information screen 1230 displays information about the provider extracted by the extraction unit 1024 of the information processing device 1000 (for example, profile information 1052 and image information 1042). For example, the provider information screen 1230 displays a profile information display area 1231, an image display area 1232 that displays multiple images 1235 to 1237 associated with the provider extracted by the extraction unit 1024 of the information processing device 1000, a button 1233 to resuggest new information, and a confirmation button 1234. Specifically, the information provision unit 1025 of the information processing device 1000 transmits screen information for displaying the provider information screen 1230 to the electronic device 1100. When the electronic device 1100 receives this screen information, the control unit 1150 displays the provider information screen 1230 on the UI unit 1160 based on the received screen information. In addition, the image display area 1232 can display other images associated with the provider based on user operation (for example, scrolling operation).
[0246] Furthermore, when a user operation is received by the reception unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on that user operation. For example, after confirming the profile displayed in the profile information display area 1231 and the multiple images 1235 to 1237 displayed in the image display area 1232, if the user decides to select the provider of the profile displayed in the profile information display area 1231 as preferred information, the user performs a selection operation (e.g., a touch operation) by selecting the select button 1234. When this select operation of the select button 1234 is performed, the control unit 1150 transmits decision information indicating that the selection operation has been performed to the information processing device 1000. When the information processing device 1000 receives the decision information, the control unit 1020 executes predetermined processing to enable some kind of exchange between the provider corresponding to the received decision information and the user. As part of this predetermined process, at least one of the following is performed: a transmission process that sends information about the user (information with the user's consent) to the provider corresponding to that decision information, and a transmission process that sends information to the user for accessing the provider corresponding to that decision information (e.g., access information such as a telephone number, email address, URL (Uniform Resource Locator), SNS (Social Networking Service)). This allows the user to easily and quickly access information that suits their preferences.
[0247] Here, after reviewing the profile displayed in the profile information display area 1231 and the multiple images 1235 to 1237 displayed in the image display area 1232, it is conceivable that the user may request further suggestions for other information. For example, a user who has reviewed the multiple images 1235 to 1237 displayed in the image display area 1232 may request further suggestions because they do not suit their preferences.
[0248] Thus, if the user requests a re-suggestion of other information, the user performs a selection operation (e.g., a touch operation) by selecting the re-suggestion button 1233. When the re-suggestion button 1233 is selected, the control unit 1150 transmits re-suggestion request information indicating that the selection operation has been performed to the information processing device 1000. When the information processing device 1000 receives this re-suggestion request information, the extraction unit 1024 extracts images from the image information DB 1040 to be newly provided to the user based on the user vector generated by the vector generation unit 1023. Here, the text information "I want a stylish mug" has been entered in the text input area 1201 of the search screen 1200 (see Figure 25). Therefore, as explained in Figure 25, it is possible to store the categories of images that have already been selected using known selection techniques in memory (e.g., the storage unit 1030), and to perform a selection process to select images to be provided to the electronic device 1100 from the image information DB 1040 based on those image categories. In other words, a selection process is executed to select a new image to be provided to the electronic device 1100 from a pre-defined category (for example, mug) based on the text information entered in the text input area 1201 of the search screen 1200. This makes it possible to provide the user with appropriate image information that meets the user's request. For example, the extraction unit 1024 compares the user vector generated by the vector generation unit 1023 with the tag information (feature quantity) stored in the feature quantity 1044 of the image information DB 1040, and extracts a new image according to the user's preference based on this comparison result. Specifically, the extraction unit 1024 calculates the difference value between each component constituting the user vector and each component constituting each tag information (each vector) for each piece of tag information, and calculates the sum of these difference values for each piece of tag information. Then, the extraction unit 1024 extracts images corresponding to a predetermined number of tag information with small sums of these difference values as new images according to the user's preference. Alternatively, the cosine similarity of each vector to be compared may be used to extract images corresponding to a predetermined number of tag information with high cosine similarity as new images according to the user's preference. Alternatively, other calculation methods may be used to extract new images that match the user's preferences.
[0249] Also, the extraction unit 1024 outputs image information (e.g., image information 1042) regarding the extracted new image to the information providing unit 1025.
[0250] [Example of selection of new image] FIG. 28 is a diagram showing a display example of a selection screen 1240 displayed on the UI unit 1160 of the electronic device 1100. The selection screen 1240 is displayed on the UI unit 1160 based on the control of the control unit 1150.
[0251] On the selection screen 1240, an image display area 1241 for displaying a plurality of images 1243 to 1250, a determination button 1251, and a re-proposal button 1252 for new information are displayed. Note that the selection screen 1240 is a display screen corresponding to the selection screen 1210 shown in FIG. 26, and the image display area 1241 and the determination button 1251 correspond to the image display area 1211 and the determination button 1221. Therefore, detailed descriptions thereof are omitted. Also, since the selection operation and the determination operation for each image are the same as the example shown in FIG. 26, detailed descriptions thereof are omitted.
[0252] Also, the re-proposal button 1252 for new information corresponds to the re-proposal button 1233 for new information shown in FIG. 27. When the selection operation of the re-proposal button 1252 for new information is performed, similar to the example shown in FIG. 27, an image (e.g., an image other than each image displayed on the selection screen 1240) to be newly provided to the user may be extracted from the image information DB 1040 based on the user vector and provided to the user, or an image extracted using other information may be provided to the user. Examples of extracting a new image using other information are shown in FIGS. 36 to 40 and the like.
[0253] Furthermore, if, after one or more images have been selected by the user, the user performs a selection operation (e.g., a touch operation) by selecting the OK button 1251, the control unit 1150 transmits image selection information regarding the one or more images selected by the user to the information processing device 1000. The tag generation process by the tag generation unit 1022, the vector generation process by the vector generation unit 1023, and the provider information extraction process by the extraction unit 1024 using this image selection information are the same as in the example described above.
[0254] In this way, it is possible to select and provide new images to the user using a user vector generated based on images selected by the user (favorite images). Furthermore, by performing each of these processes multiple times, it is possible to refine the recommendations to the user. Alternatively, new images may be selected using past calculation results (for example, user vectors from one or several times ago). For example, it is possible to sequentially add up past calculation results and use the value obtained by dividing by the number of calculation results being added.
[0255] [Example of server operation] Figure 29 is a flowchart illustrating an example of a selection process performed by the information processing device 1000. This selection process is executed by the control unit 1020 (see Figure 21) based on a program stored in the memory unit 1030 (see Figure 21). This selection process is executed, for example, when an image request is received from the electronic device 1100. This selection process will be explained with reference to Figures 1 to 28 as appropriate.
[0256] In step S1301, the information provision unit 1025 transmits image information to the electronic device 1100 based on an image request from the electronic device 1100. This image information is information for displaying the selection screen 1210 (see Figure 26) on the UI unit 1160 of the electronic device 1100. For example, the electronic device 1100 may receive the services in this embodiment by installing a predetermined application (search application) on the electronic device 1100. In this case, the control unit 1150 of the electronic device 1100 transmits an image request to the information processing device 1000 by launching the search application on the electronic device 1100. Alternatively, as shown in Figure 25, the selection screen 1210 (see Figure 26) may be displayed in response to user input. In this case, the electronic device 1100 transmits an image request containing the input information (e.g., text information) to the information processing device 1000. The extraction unit 1024 then selects an image based on the information (e.g., text information) contained in the image request. A known image search method can be used for this selection process.
[0257] In step S1302, the information acquisition unit 1021 determines whether or not it has received image selection information (preferred image information) from the electronic device 1100. For example, if the user selects an image on the selection screen 1210 (see Figure 26) and then selects the OK button 1221, the control unit 1150 of the electronic device 1100 transmits the image selection information to the information processing device 1000. If the image selection information is received, the process proceeds to step S1303. On the other hand, if the image selection information is not received, monitoring continues.
[0258] In step S1303, the tag generation unit 1022 retrieves tag information associated with one or more images corresponding to the image selection information received in step S1302 from the image information DB 1040. The method for retrieving this tag information is the same as the method described above.
[0259] In step S1304, the vector generation unit 1023 generates a user vector relating to the user of the electronic device 1100 that transmitted the image selection information, based on the tag information acquired in step S1303. The method for generating this user vector is the same as the generation method described above.
[0260] In step S1305, the extraction unit 1024 extracts and selects a provider from among the providers stored in the provider information DB 1050 that matches the user's preferences corresponding to the user vector, based on the user vector generated in step S1304. The method for selecting this provider (extraction method) is the same as the extraction method described above.
[0261] In step S1306, the information provision unit 1025 transmits provider information (for example, profile information 1052, image information 1042) related to the provider selected in step S1305 to the electronic device 1100. When the electronic device 1100 receives this provider information, the control unit 1150 of the electronic device 1100 displays the provider information screen 1230 (see Figure 27) on the UI unit 1160.
[0262] In step S1307, the information acquisition unit 1021 determines whether or not it has received decision information from the electronic device 1100 that transmitted the provider information in step S1306. For example, if the user selects the decision button 1234 on the provider information screen 1230 (see Figure 27), the control unit 1150 of the electronic device 1100 transmits decision information to the information processing device 1000. If decision information is received, the process proceeds to step S1308. On the other hand, if decision information is not received, the process proceeds to step S1309.
[0263] In step S1308, the control unit 1020 performs a predetermined process to introduce the provider (selection provider) corresponding to the decision information received in step S1307 to the user of the electronic device 1100. This predetermined process is the same as the predetermined process described above.
[0264] In step S1309, the information acquisition unit 1021 determines whether or not it has received a request for resubmission information from the electronic device 1100 that transmitted the provider information in step S1306. For example, if the user selects the resubmission button 1233 on the provider information screen 1230 (see Figure 27), the control unit 1150 of the electronic device 1100 transmits the request for resubmission information to the information processing device 1000. If the request for resubmission information is received, the process proceeds to step S1310. On the other hand, if the request for resubmission information is not received, the process returns to step S1307.
[0265] In step S1310, the extraction unit 1024 extracts images from the image information DB 1040 to be newly provided to the user based on the user vector generated in step S1304. This extraction process is the same as the extraction process described above (for example, the extraction process shown in Figure 27).
[0266] In step S1311, the information provision unit 1025 transmits image information (for example, image information 1042) related to the new image extracted in step S1311 to the electronic device 1100. This image information is used to display the selection screen 1240 (see Figure 28) on the UI unit 1160 of the electronic device 1100.
[0267] [Example of effect] Here, it is assumed that while users have an abstract ideal image or preference regarding the information they desire, it is difficult for them to clearly articulate that image or desire. Therefore, if the user is asked to answer questions about the image or desire of the information they desire, and information is suggested based on these answers, there is a risk that appropriate information may not be suggested. In this embodiment, however, it is possible to appropriately acquire the user's preferences by selecting an image according to the user's preferences from among multiple images displayed on the electronic device 1100. Then, a user vector is generated using tag information related to the image according to the user's preferences, and information (provider, image) is suggested to the user based on the comparison result between this user vector and the provider vector (or image vector), making it possible to suggest appropriate information (provider, image) according to the user's preferences. In other words, it is possible to support the acquisition of desired information in an intuitive and easy-to-understand manner.
[0268] [Example of user-edited image] It is conceivable that among the images provided to the user, none may match the user's preferences. In this case, it may be difficult for the user to select an image they like. Therefore, it is conceivable to allow the user to temporarily edit the images provided to them to suit their preferences, thereby making it easier for the user to select an image they like.
[0269] [Example of Item Information Database Configuration] Figure 30 is a simplified diagram showing the contents of the item information stored in the item information DB 1070. Note that the item information DB 1070 may be stored in the storage unit 1030 of the information processing device 1000, or it may be stored and used in an external device other than the information processing device 1000.
[0270] Item Information DB1070 is a database for storing items used to edit images provided by providers, and associating these items with the feature quantities of multiple tags generated for each item.
[0271] Specifically, image information 1072, tag 1073, and feature quantity 1074 are associated with item identification information 1071 and stored in item information DB 1070.
[0272] Item identification information 1071 is identification information used to identify each item prepared by the service provider. Here, the service provider refers to a business operator that operates the information processing system IS1 and provides various services to users. In Figure 30, for the sake of simplicity, an example is shown in which only the name indicating the item is stored as identification information in item identification information 1071, but other types of identification information may also be used.
[0273] Image information 1072 is information about an item image prepared by the service provider so that the user can edit it. In Figure 30, for the sake of simplicity, an example is shown in which only the image is stored in image information 1072, but various attribute information (e.g., size information) associated with the image may also be included.
[0274] Tag 1073 is information indicating the tag generated for the item image stored in image information 1072. Feature 1074 is a feature assigned to the tag stored in tag 1073. This tag information is used when generating the user vector. Note that tag information with the same content as feature 1044 shown in Figure 23 may be stored, or tag information with different content from feature 1044 shown in Figure 23 may be stored. Furthermore, only the necessary information from these may be stored, and other information may be omitted.
[0275] [Example of a user-edited image screen] Figure 31 shows an example of the display of the editing screen 1260 shown on the UI unit 1160 of the electronic device 1100. The editing screen 1260 is displayed on the UI unit 1160 based on the control of the control unit 1150.
[0276] The editing screen 1260 displays an image display area 1261 for displaying multiple images 1262 to 1265, a confirmation button 1266, and an item display area 1270. Specifically, the information provision unit 1025 of the information processing device 1000 transmits editing screen information for displaying the editing screen 1260 to the electronic device 1100. When the electronic device 1100 receives this editing screen information, the control unit 1150 displays the editing screen 1260 on the UI unit 1160 based on the received editing screen information. The image display area 1261 corresponds to the image display area 1211 shown in Figure 26.
[0277] The item display area 1270 is an area that displays each item image 1271 to 1276 whose item information is stored in the item information DB 1070 (see Figure 30). The user can generate an image of their choice by performing a move operation to move the desired item displayed in the item display area 1270 to the desired image displayed in the image display area 1261. For example, if the UI unit 1160 is configured as a touch panel, the user can perform an operation to move the desired item to the desired position on the image while touching it (e.g., a drag-and-drop operation). For example, the user can place the item image 1272, which represents a cactus, to the desired position on image 1262 by performing an operation to move it to the desired position on image 1262. This movement transition is schematically shown by arrow 1277. Furthermore, the image on which the item image is placed (edited image) is in a selected state. Furthermore, known image processing can be used for the image synthesis process that places and combines the item image within the image. Furthermore, each item image may be made capable of various image processing operations (e.g., scaling, 3D angle adjustment, image addition / editing) based on user input. Known image processing techniques can be used for these operations. Additionally, item images placed within an image may be appropriately transformed in relation to other elements in the image using known image processing techniques. For example, item images can be scaled based on the size of each object in the image, or their angles can be adjusted based on the tilt of each object in the image.
[0278] Thus, if no image matching the user's preferences is displayed in the image display area 1261, it is possible to edit a portion of each image displayed in the image display area 1261 to generate an image that matches the user's preferences. Note that the image generated through this editing is a temporary image generated for the purpose of selecting a provider or a new image, and is deleted after the process of selecting a provider or a new image is completed.
[0279] In this manner, when a user operation is received by the reception unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on that user operation. For example, when a user performs a selection operation (e.g., a touch operation) to select an image to be edited from among the images displayed in the image display area 1261, the control unit 1150 sets the preferred image selected by that operation to a selected state.
[0280] Furthermore, if the user edits one or more images and then performs a selection operation (for example, a touch operation) by selecting the OK button 1266, the control unit 1150 transmits edited image information relating to the one or more images edited by the user to the information processing device 1000. When the information processing device 1000 receives this edited image information, the information acquisition unit 1021 outputs the received edited image information to the tag generation unit 1022 and the vector generation unit 1023.
[0281] [Example of user vector generation] The tag generation unit 1022 obtains tag information from the image information DB 1040 and the item information DB 1070 for one or more edited images corresponding to the edited image information transmitted from the electronic device 1100. Specifically, the tag generation unit 1022 obtains the tag information associated with the original image of each edited image from the image information DB 1040, and obtains the tag information associated with the items added to each edited image through editing from the item information DB 1070. The tag generation unit 1022 then outputs the obtained tag information to the vector generation unit 1023.
[0282] Next, the vector generation unit 1023 generates a user vector related to the user who sent the edited image information, based on the tag information output from the tag generation unit 1022. In this case, the method is the same as the user vector generation method described above, except that the tag information associated with the item is used, so a detailed explanation is omitted here.
[0283] Here, it is assumed that edited images, to which items have been added by the user, reflect the user's preferences more than unedited images. Therefore, tag information associated with edited images to which items have been added by the user may be weighted more heavily than other tag information when generating user vectors. It is also conceivable that edited images may have multiple items added. In the case of such edited images, it is assumed that the user's preferences are reflected even more. Therefore, for example, the weighting of the edited image and the weighting of the added items may be increased according to the number of items added to the edited image when generating user vectors. In other words, the weighting of the edited image and the weighting of the added items may be changed based on various information about the items added to the edited image (e.g., the number of items, the proportion of item images in the image) when generating user vectors. This makes it possible to generate user vectors that reflect the user's preferences.
[0284] Here, as shown in Figure 31, when an image is edited by adding items to it, it is conceivable that the design style of the object (for example, a mug) may differ depending on the position of the item within that object. Therefore, the tag generation unit 1022 may generate new tag information for the edited image using the AI model described above. In this case, it becomes possible to generate more appropriate tag information according to the editing results made by the user.
[0285] [Examples of electronic device operation] Figure 32 is a flowchart illustrating an example of image editing processing by the electronic device 1100. This image editing processing is executed by the control unit 1150 (see Figure 21) based on a program stored in the memory unit 1170 (see Figure 21). This image editing processing is executed, for example, when a user operation is performed to request an editing screen. This image editing processing will be explained with reference to Figures 1 to 31 as appropriate.
[0286] In step S1321, the control unit 1150 sends editing screen request information to the information processing device 1000 to request editing screen information. For example, when a user operation is performed to request an editing screen, the control unit 1150 sends editing screen request information to the information processing device 1000.
[0287] In step S1322, the control unit 1150 determines whether or not it has received editing screen information from the information processing device 1000. If editing screen information is received, the process proceeds to step S1323. On the other hand, if editing screen information is not received, monitoring continues.
[0288] In step S1323, the control unit 1150 displays an editing screen (for example, editing screen 1260 (see Figure 31)) on the UI unit 1160 based on the editing screen information received in step S1322.
[0289] In step S1324, the control unit 1150 determines whether or not a user operation has been received by the reception unit 1161 regarding the editing screen displayed in step S1323. If a user operation has been received, the process proceeds to step S1325. On the other hand, if no user operation has been received, monitoring continues.
[0290] In step S1325, the control unit 1150 determines whether the user operation received in step S1324 is an editing operation for editing an image displayed on the editing screen. For example, if a move operation is performed to move a desired item displayed in the item display area 1270 (see Figure 31) to a desired image displayed in the image display area 1261 (see Figure 31), it is determined that an editing operation has been performed as the user operation. If the user operation is an editing operation, the process proceeds to step S1326. On the other hand, if the user operation is not an editing operation, the process proceeds to step S1327.
[0291] In step S1326, the control unit 1150 executes an editing process to edit the image based on the user operation (editing operation) received in step S1324. For example, if a move operation is performed to move an item displayed in the item display area 1270 to an image displayed in the image display area 1261, image processing is performed to move the item and combine it with the image in accordance with that move operation. Also, for example, if a delete operation is performed to delete an item that has been moved to the image, image processing is performed to delete the item in accordance with that delete operation.
[0292] In step S1327, the control unit 1150 determines whether the user operation received in step S1324 is a decision operation performed after one or more images have been edited. For example, if a selection operation (e.g., a touch operation) is performed to select the decision button 1266 (see Figure 31), it is determined that a decision operation has been performed as the user operation. If the user operation is a decision operation, the process proceeds to step S1329. On the other hand, if the user operation is not a decision operation, the process proceeds to step S1328.
[0293] In step S1328, the control unit 1150 executes a predetermined process based on the user operation received in step S1324. For example, if a selection operation is performed to select an image displayed in the image display area 1261, image processing is performed to select the image in accordance with that selection operation.
[0294] In step S1329, the control unit 1150 transmits edited image information relating to one or more images edited by the user to the information processing device 1000.
[0295] In this way, by allowing users to edit images from among multiple images displayed on the electronic device 1100 according to their preferences, it becomes possible to express the user's preferences more appropriately based on the user's active editing. Therefore, it becomes possible to suggest appropriate information (providers, images) according to the user's preferences based on the user's editing results. In other words, it becomes possible to support the acquisition of desired information in an intuitive and easy-to-understand manner.
[0296] The above examples illustrate editing images using pre-prepared images (item images), but images may also be edited using text information entered by the user (e.g., manual input, voice input, gesture input). For example, the user might enter "I like piano" or "I like sofas" as text information. In this case, it is possible to reflect the entered text information in the selected image. For example, it is possible to generate an image related to the entered text information using a generative AI that has learned about each tag (e.g., an image generation AI (e.g., Stable Diffusion, Midjourney)). Then, the desired edited image can be generated by combining the generated image (the image related to the text information) with the selected image. In this case, for example, known image synthesis techniques can be used. This makes it possible to generate the desired edited image related to the entered text information. In this case, for example, the weight associated with the edited image and the weight associated with the entered text information may be increased according to the weight of the entered text information (e.g., the number of characters) to generate a user vector. This makes it possible to generate a user vector that reflects the user's preferences.
[0297] [Example using images obtained by the user] As mentioned above, it is conceivable that none of the images provided to the user may match their preferences. In this case, it may be difficult for the user to select their preferred image. Therefore, it is conceivable that information tailored to the user's preferences can be selected by generating a user vector based on the images acquired by the user.
[0298] [Example of image information database configuration] Figure 33 is a simplified diagram showing the contents of the image information stored in the image information DB1080.
[0299] The image information DB 1080 is a database stored in the storage unit 1170 of the electronic device 1100 (see Figure 21), or in an external device (for example, a server). The image information DB 1080 also stores image information acquired by the image acquisition unit 1130. For example, if a smartphone is used as the electronic device 1100, the camera installed in the smartphone corresponds to the image acquisition unit 1130.
[0300] When the user performs an image acquisition operation (for example, taking a still image or taking a video), the control unit 1150 of the electronic device 1100 records the image acquired by the image acquisition unit 1130 based on that operation as an image file (still image file, video file) in the image information DB 1080. For example, as shown in Figure 33, the image file is stored in image information 1082 in association with image identification information 1081.
[0301] In this case, the control unit 1150 stores the location information (e.g., latitude and longitude) acquired by the location information acquisition unit 1120 of the electronic device 1100, associating it with the image file, based on the timing of the acquisition operation. For example, as shown in Figure 33, the acquired location information is stored in location information 1083 as associated with the image file. For example, if a smartphone is used as the electronic device 1100, the GPS device installed in the smartphone corresponds to the location information acquisition unit 1120.
[0302] Furthermore, the control unit 1150 stores the audio information acquired by the sound acquisition unit 1140 of the electronic device 1100 in association with the image file, based on the timing of the acquisition operation. For example, as shown in Figure 33, the acquired audio information is stored in audio information 1084 in association with the image file. For example, if a smartphone is used as the electronic device 1100, the microphone installed in the smartphone corresponds to the sound acquisition unit 1140. When taking a still image, it is possible to store audio information for a predetermined time based on the timing of the capture in association with the still image. Also, when shooting a video, the acquired audio information may be used to select a representative image from the video (for example, the first frame, the frame with the most distinctive features) and use it as a still image.
[0303] For example, a user can use the electronic device 1100 to photograph various subjects, landscapes, etc. that suit their preferences, and images that suit the user's preferences will be stored in the image information DB 1080. Therefore, it is possible to generate a user vector according to the user's preferences using the image information stored in the image information DB 1080. Note that if a stationary device such as a personal computer is used as the electronic device 1100, it may not include the location information acquisition unit 1120 and the image acquisition unit 1130. In this case, external devices may be used as the location information acquisition unit 1120 and the image acquisition unit 1130, and image information acquired using an imaging device (e.g., a smartphone, digital still camera, digital video camera (e.g., a camera-integrated recorder)) can be acquired, stored in the image information DB 1080, and used.
[0304] [Example of user-selected images] Figure 34 shows an example of the display of the selection screen 1280 shown on the UI unit 1160 of the electronic device 1100. The selection screen 1280 is displayed on the UI unit 1160 based on the control of the control unit 1150.
[0305] The selection screen 1280 displays an image display area 1290 for displaying multiple images 1291 to 1296, a selected image display area 1281 for displaying the selected image, a location information specification area 1283, an audio information specification area 1284, and a confirmation button 1285. Specifically, the information provision unit 1025 of the information processing device 1000 transmits selection screen information for displaying the selection screen 1280 to the electronic device 1100. When the electronic device 1100 receives this selection screen information, the control unit 1150 displays the selection screen 1280 on the UI unit 1160 based on the received selection screen information.
[0306] The image display area 1290 is an area for displaying each of the images 1291 to 1296 whose image information is stored in the image information DB 1080 of the electronic device 1100. In other words, the control unit 1150 acquires each of the image information stored in the image information DB 1080 and displays it in the image display area 1290 of the selection screen 1280. There are no particular limitations on the images to be displayed here, and various types of images can be displayed.
[0307] The user can select their preferred image by performing a move operation to move the desired image displayed in the image display area 1290 to the selected image display area 1281. For example, if the UI unit 1160 is configured as a touch panel, the user can perform an operation (e.g., drag and drop) to move the desired image to the selected image display area 1281 while touching it. For example, by performing an operation to move the image 1292 containing a mug to the selected image display area 1281, the user can place the image 1292 in the selected image display area 1281. This movement transition is schematically shown by arrow 1297. Note that an image already displayed in the selected image display area 1281 and other images may be combined using known image processing techniques to create a composite image. Furthermore, one or more images may be edited using the editing techniques described above.
[0308] The selected image display area 1281 is an area that displays each of the images 1218 and 1292 selected based on user operation. The selected image display area 1281 may also display each of the images selected in the screens shown in Figures 26 and 28, and each of the edited images edited in the editing screen 1260 shown in Figure 31. When various images are displayed in this way, the user may be able to perform deletion, addition, and editing operations on various images as appropriate.
[0309] Figure 34 shows an example of displaying the image 1218, which was selected in Figure 26, in the selected image display area 1281.
[0310] The location information specification area 1283 is an area operated on when specifying whether or not to use the location information associated with the image displayed in the image display area 1290. For example, after selecting an image moved to the selected image display area 1281, if a selection operation is performed by selecting the pull-down button in the location information specification area 1283, a selection area for selecting whether to use or not will be displayed, and the user will perform the desired operation in this selection area. For example, if it is specified to use the location information associated with the selected image, it is possible to generate a user vector using the location information associated with that image. On the other hand, if it is specified not to use the location information associated with the selected image, it is possible to generate a user vector without using the location information associated with that image.
[0311] The audio information specification area 1284 is an area operated on when specifying whether or not to use the audio information associated with the image displayed in the image display area 1290. For example, after selecting an image moved to the selected image display area 1281, if a selection operation is performed by selecting the pull-down button in the audio information specification area 1284, a selection area for selecting whether to use or not will be displayed, and the user will perform the desired operation in this selection area. For example, if it is specified to use the audio information associated with the selected image, it is possible to generate a user vector using the audio information associated with that image. On the other hand, if it is specified not to use the audio information associated with the selected image, it is possible to generate a user vector without using the audio information associated with that image.
[0312] In this manner, when a user operation is received by the reception unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on that user operation. Furthermore, if the user performs a selection operation (for example, a touch operation) by selecting the OK button 1285 after moving one or more images to the selected image display area 1281, the control unit 1150 transmits selected image information regarding the one or more images selected by the user to the information processing device 1000. This selected image information includes image information, location information (if specified for use), and audio information (if specified for use). When the information processing device 1000 receives this selected image information, the information acquisition unit 1021 outputs the received selected image information to the tag generation unit 1022 and the vector generation unit 1023.
[0313] [Example of user vector generation] The tag generation unit 1022 generates tag information for one or more selected images corresponding to the selected image information transmitted from the electronic device 1100. If an image stored in the image information DB 1040 (see Figure 23) is selected as the image, the tag generation unit 1022 retrieves the tag information for that image from the image information DB 1040.
[0314] Furthermore, if the image stored in the image information DB 1080 of the electronic device 1100 is a selected image, the tag generation unit 1022 generates tag information for that image. For example, it is possible to generate tag information for that image using the AI model described above.
[0315] Furthermore, if at least one of location information (if specified for use) and audio information (if specified for use) is included in the selected image information, tag information can be generated for each of these pieces of information. For example, a database associating location information (e.g., latitude and longitude) with tag information (e.g., information pre-set for each region) can be prepared in advance in the storage unit 1030 of the information processing device 1000, and tag information corresponding to the location information included in the selected image information can be obtained using this database.
[0316] Furthermore, for example, a database linking location information (e.g., latitude and longitude) with the characteristics and environmental image of a certain place (e.g., Okinawa is "hot," Kamakura is "ancient capital," "calm") can be pre-prepared in the storage unit 1030 of the information processing device 1000. Using this database, it is possible to acquire the characteristics and environmental image corresponding to the location information contained in the selected image information (e.g., Okinawa is "hot," Kamakura is "ancient capital," "calm") and generate text information related to this characteristic and environmental image (e.g., "hot," "ancient capital," "calm"). In this case, it is possible to generate tag information from the text information using the AI model described later.
[0317] Furthermore, for example, a database is prepared in advance in the storage unit 1030 of the information processing device 1000 that associates audio information with tag information (for example, information pre-set for each type of sound (e.g., "noise," "quiet," "babbling brook," "children's voices")), and this database can be used to obtain tag information corresponding to the audio information contained in the selected image information.
[0318] Furthermore, it is possible to convert audio information into text information, for example. In this case, it is possible to generate tag information from the text information using the AI model described later. Alternatively, tag information may be generated using an AI model trained with location information, audio information, etc.
[0319] Next, the vector generation unit 1023 generates a user vector related to the user who sent the image selection information, based on the tag information output from the tag generation unit 1022. The method for generating this user vector is the same as the method described above, so a detailed explanation is omitted here.
[0320] The above example demonstrates generating tag information using a pre-trained AI model and then generating user vectors using this tag information. However, other AI models may be used. For example, a large language model (LLM) can be used as the AI model. Examples of LLMs that can be used include ChatGPT (Generative Pre-trained Transformer), Bard, Llama2 (Large Language Model Meta AI 2), Gemini, and Claude. These are just examples, and other AI models may be used. Furthermore, images (e.g., advertising images) may also have accompanying text information. In such cases, the AI model can be used to generate tag information related to that text information, and this tag information, combined with the tag information related to the image (e.g., advertising image), can be used to generate user vectors.
[0321] [Examples of electronic device operation] Figure 35 is a flowchart illustrating an example of image selection processing by the electronic device 1100. This image selection processing is executed by the control unit 1150 (see Figure 21) based on a program stored in the memory unit 1170 (see Figure 21). This image selection processing is executed, for example, when a user operation is performed to request a selection screen. This image selection processing will be explained with reference to Figures 1 to 34 as appropriate.
[0322] In step S1331, the control unit 1150 transmits selection screen request information to the information processing device 1000 to request selection screen information. For example, when a user operation is performed to request a selection screen, the control unit 1150 transmits selection screen request information to the information processing device 1000.
[0323] In step S1332, the control unit 1150 determines whether or not it has received selection screen information from the information processing device 1000. If it has received selection screen information, it proceeds to step S1333. On the other hand, if it has not received selection screen information, it continues monitoring.
[0324] In step S1333, the control unit 1150 obtains image information (for example, image information 1082) to be displayed on the editing screen corresponding to the selection screen information received in step S1332 from the image information DB 1080 (see Figure 33).
[0325] In step S1334, the control unit 1150 displays a selection screen (for example, selection screen 1280 (see Figure 34)) on the UI unit 1160 based on the selection screen information received in step S1332. This selection screen displays the image information acquired in step S1333.
[0326] In step S1335, the control unit 1150 determines whether or not a user operation has been received by the reception unit 1161 regarding the selection screen displayed in step S1334. If a user operation has been received, the process proceeds to step S1336. On the other hand, if no user operation has been received, monitoring continues.
[0327] In step S1336, the control unit 1150 determines whether the user operation received in step S1335 is a selection operation to select an image displayed on the selection screen. For example, if a move operation is performed to move an image displayed in the image display area 1290 (see Figure 34) to the selected image display area 1281 (see Figure 34), it is determined that a selection operation has been performed as the user operation. If the user operation is a selection operation, the process proceeds to step S1337. On the other hand, if the user operation is not a selection operation, the process proceeds to step S1338.
[0328] In step S1337, the control unit 1150 executes a selection process to select an image based on the user operation (selection operation) received in step S1335. For example, if a move operation is performed to move an image displayed in the image display area 1290 (see Figure 34) to the selected image display area 1281 (see Figure 34), image processing is performed to move the image to the selected image in accordance with that move operation. Also, for example, if a delete operation is performed to delete an image that has been moved to the selected image display area 1281 as a selected image, image processing is performed to delete the image corresponding to that delete operation from the selected image display area 1281.
[0329] In step S1338, the control unit 1150 determines whether the user operation received in step S1335 is a decision operation performed after one or more images have been selected. For example, if a selection operation (e.g., a touch operation) is performed to select the decision button 1285 (see Figure 34), it is determined that a decision operation has been performed as the user operation. If the user operation is a decision operation, the process proceeds to step S1340. On the other hand, if the user operation is not a decision operation, the process proceeds to step S1339.
[0330] In step S1339, the control unit 1150 executes a predetermined process based on the user operation received in step S1335. For example, if a selection operation is performed to select a pull-down button in the location information specification area 1283 or the voice information specification area 1284 (see Figure 34), image processing is performed to use or not use the corresponding area according to that selection operation.
[0331] In step S1340, the control unit 1150 transmits selected image information relating to one or more images selected by the user to the information processing device 1000.
[0332] In this way, by allowing users to select images according to their preferences from those acquired using the electronic device 1100, it becomes possible to more appropriately express user preferences based on the user's proactive image collection. Therefore, it becomes possible to suggest appropriate information (providers, images) according to the user's preferences based on the user's image collection behavior (e.g., shooting behavior). In other words, it becomes possible to support the acquisition of desired information in an intuitive and easy-to-understand manner.
[0333] [Example using text information] The above examples primarily used images to select providers and new images. However, it is conceivable that users may not be presented with providers or new images that match their preferences. Therefore, the following examples demonstrate how to select providers and new images using text information.
[0334] [Example of generating user vectors based on text information] Figure 36 shows an example of the display of the text input screen 1400 shown on the UI unit 1160 of the electronic device 1100. The text input screen 1400 is displayed on the UI unit 1160 based on the control of the control unit 1150.
[0335] The text input screen 1400 displays an image display area 1401 for displaying multiple images 1402 to 1404, a text information input field 1405, a suggestion button 1406, and a confirmation button 907. Specifically, the information provision unit 1025 of the information processing device 1000 transmits screen information for displaying the text input screen 1400 to the electronic device 1100 based on a request from the electronic device 1100. When the electronic device 1100 receives this screen information, the control unit 1150 displays the text input screen 1400 on the UI unit 1160 based on the received screen information. It is also possible to display other images on the text input screen 1400 based on user operation.
[0336] Furthermore, when user input is received by the reception unit 1161 of the UI unit 1160, the control unit 1150 executes various processes based on that user input. Here, it is conceivable that the user's preferred image may not be displayed in the image display area 1401. In such cases, the user can input text related to their preferred image (i.e., the product or service in the image) into the text information input field 1405 using user input (e.g., manual operation, voice input). Figure 36 shows an example where the user inputs "I prefer a more modern design." Note that this is just one example, and other text related to the preferred image may be entered, such as "I like something a little simpler," "I like a cuter design," or "I want an ethnic nuance."
[0337] Furthermore, if, after entering text in the text information input field 1405, the user performs a selection operation (e.g., a touch operation) by selecting the re-suggestion button 1406, the control unit 1150 transmits text information corresponding to the characters entered by the user to the information processing device 1000. When the information processing device 1000 receives this text information, the information acquisition unit 1021 outputs the received text information to the tag generation unit 1022 and the vector generation unit 1023. If one or more images are selected by the user, the control unit 1150 transmits image selection information regarding the selected one or more images along with the text information to the information processing device 1000. In this case, categories (e.g., mugs) that have been narrowed down in advance based on the text information (e.g., I want a stylish mug) entered in the text input area 1201 of the initial search screen (e.g., search screen 1200 (see Figure 25)) are stored in the memory of the information processing device 1000. Therefore, a selection process is executed to select new images to provide to the electronic device 1100 based on the categories stored in that memory.
[0338] [Example of user vector generation] The tag generation unit 1022 generates tag information corresponding to the text information transmitted from the electronic device 1100. If one or more images are selected by the user, image selection information regarding those selected images is transmitted from the electronic device 1100 along with the text information. In this case, it is possible to generate a user vector considering this image selection information. An example of this is shown in Figures 39 and 40.
[0339] For example, the tag generation unit 1022 can generate tag information corresponding to text information using a predetermined database (e.g., a dictionary database) for converting text information into tag information. For example, it is possible to use a dictionary database that associates strings with the characteristics of each element constituting the tag information (e.g., each feature corresponding to tag 1043 shown in Figure 23). The tag generation unit 1022 then extracts strings from the dictionary database from among the strings contained in the text corresponding to the text information transmitted from the electronic device 1100, and extracts the characteristics of the tag information corresponding to the extracted strings from the dictionary database. Known character recognition technology can be used as the method for extracting strings. If multiple strings are extracted, the characteristics of the tag information corresponding to each string are extracted. The tag generation unit 1022 then aggregates the characteristics of one or more extracted tag information (e.g., calculates the average value) to obtain tag information corresponding to the text information transmitted from the electronic device 1100. The tag generation unit 1022 then outputs the generated tag information to the vector generation unit 1023.
[0340] Furthermore, it is possible to generate tag information (user vectors) from text information using an AI model. For example, it is possible to use an AI model that has learned multiple text pieces of information to which predetermined tags, which associate sentences with the characteristics of each element constituting the tag information (for example, each feature corresponding to tag 1043 shown in Figure 23), are assigned as training labels. The learning method shown in Figure 22 can be applied to this learning method.
[0341] Furthermore, it is possible to use an AI model such as LLM. Examples of LLMs that can be used include ChatGPT, Bard, Llama2, Gemini, and Claude. Note that these are just examples, and other AI models may be used. For example, conditional information that specifies that the LLM should output image features for each of multiple tags based on text information entered by the user can be input as instruction information (prompt), and the output information in response to this can be used as tag information.
[0342] Here, we show an example of a prompt that suggests tags using ChatGPT. For example, it is possible to input instruction information (prompt) to the LLM that uses text information sent from the electronic device 1100 to output a score (0 to 1) regarding its association with each of several tags (for example, tag 1043 shown in Figure 23), and to use the output information as feature quantities for each of the multiple tags. Note that the tags are not limited to those shown in Figure 23, and other tags may be used. For example, tags such as Western, Japanese, modern, American, antique, country, simple, natural, Scandinavian, industrial, woody, family gathering / warmth, high-class / elegant / luxury, cool / chic, casual / pop, sophisticated / stylish, openness, calming, flashy, substantial, cute, etc. may be used as elements of tags.
[0343] Next, the vector generation unit 1023 generates a user vector relating to the user who sent the text information, based on the tag information output from the tag generation unit 1022. In this case, the tag information generated by the tag generation unit 1022 can be used as the user vector. If there are multiple pieces of text information transmitted from the electronic device 1100, tag information is generated for each piece of text information, so it is possible to generate a user vector using multiple pieces of tag information, similar to the example shown in Figure 26.
[0344] [Example of image selection based on text information] Next, the extraction unit 1024 extracts images from the image information DB 1040 to be newly provided to the user based on the user vectors generated by the vector generation unit 1023. This extraction method is the same as the method for extracting new images when the new information re-suggestion button 1233 (see Figure 27) is selected, so a detailed explanation is omitted here. For example, the extraction process for new images is performed using a category (e.g., mug) held in the memory of the information processing device 1000.
[0345] Furthermore, the extraction unit 1024 outputs image information (for example, image information 1042) related to the extracted image to the information provision unit 1025. The information provision unit 1025 then transmits the image information to the electronic device 1100, which displays the image information.
[0346] [Example of selecting a new image] Figure 37 shows an example of the display of the selection screen 1410 shown on the UI unit 1160 of the electronic device 1100. The selection screen 1410 is displayed on the UI unit 1160 based on the control of the control unit 1150. The selection screen 1410 is a modified version of the selection screen 1240 shown in Figure 28, with a modified portion of the upper message; the rest is the same as the selection screen 1240. Therefore, the parts common to the selection screen 1240 are denoted by the same reference numerals. The image display area 1241 displays each image selected by the image selection process based on the text information described above, but here, for the sake of simplicity, an example is shown in which multiple images 1243 to 1250, similar to those on the selection screen 1240, are displayed.
[0347] In this way, it is possible to select and provide images to the user using a user vector generated based on text information entered by the user. Furthermore, a text information input field 1405 (see Figure 36) may be provided on the selection screen 1410 to allow for additional text information input. In this case, when text information is entered by the user, it is possible to propose even newer images to the user based on that text information. Moreover, by performing each of these processes multiple times, it is possible to refine the recommendations to the user.
[0348] The above examples illustrate how to accept text information via user input when the user does not find a preferred image among those presented, or when the user has requests for specific images. However, the system is not limited to these examples. For instance, similar to the example shown in Figure 25, it is possible to enable the acceptance of text information via user input before images are presented, and then use a user vector generated based on this text information to select and present images to the user.
[0349] [Example of server operation] Figure 38 is a flowchart showing an example of a selection process performed by the information processing device 1000. This selection process is a modified version of the selection process shown in Figure 29. Specifically, steps S1351 to S1354 have been added. Except for this addition, the process is the same as the selection process shown in Figure 29, so the same reference numerals are used for parts common to both Figure 29 and Figure 38, and their explanations are omitted.
[0350] In step S1351, the information acquisition unit 1021 determines whether or not it has received text information from the electronic device 1100. For example, if the user inputs text information on the text input screen 1400 (see Figure 36) and then selects the re-suggestion button 1406, the control unit 1150 of the electronic device 1100 transmits the text information to the information processing device 1000. If text information is received, the process proceeds to step S1352. On the other hand, if text information is not received, the process returns to step S1302.
[0351] In step S1352, the tag generation unit 1022 and the vector generation unit 1023 generate tag information (user vectors) based on the text information received in step S1351. The method for generating these user vectors is the same as the generation method described above.
[0352] In step S1353, the extraction unit 1024 extracts images from the image information DB 1040 to be newly provided to the user based on the user vector generated in step S1352.
[0353] In step S1354, the information provision unit 1025 transmits image information (for example, image information 1042) related to the new image extracted in step S1353 to the electronic device 1100. This image information is used to display the selection screen 1410 (see Figure 37) on the UI unit 1160 of the electronic device 1100.
[0354] [Example of image re-recommendation based on image and text information] The above example demonstrated how to present images to a user using user vectors generated based on text information. It is also conceivable that after presenting multiple images to the user, the user will select their preferred image and provide text-based requests regarding that image. Therefore, the following example demonstrates how to generate new user vectors based on the images selected by the user and the user's requests regarding the images presented.
[0355] [Example of generating a user vector based on selected image and text information] Figure 39 shows an example of the display of the text input screen 1430 shown on the UI unit 1160 of the electronic device 1100. The text input screen 1430 is the same as the text input screen 1400 shown in Figure 36, except that the content entered in the text information input field 1405 and the images 1402 and 1403 are selected (star area is added). Otherwise, it is the same as the text input screen 1400. Therefore, the parts that are common with the text input screen 1400 are indicated with the same reference numerals. The transmission process of each piece of information to the information processing device 1000 is also the same as the example shown in Figure 36.
[0356] As mentioned above, it is conceivable that the image display area 1401 may not display the user's preferred image, that there may be few images of the user's preference, or that an image slightly different from the user's preference may be displayed. In such cases, the user can input text related to the preferred image into the text information input field 1405 using user input (e.g., manual operation or voice input). Figure 39 shows an example where the user inputs "Please recommend a mug with a brighter atmosphere." Note that this is just one example, and the user could also input text related to the preferred image, such as "Please clear the original recommendation results and recommend an image of a mug with a brighter atmosphere."
[0357] [Example of user vector generation] The tag generation unit 1022 generates tag information corresponding to the image selection information and text information transmitted from the electronic device 1100. The generation of tag information based on image selection information is the same as the example shown in Figure 26, etc. The generation of tag information based on text information is the same as the example shown in Figure 36. Furthermore, the generation of user vectors (user vectors based on image selection information and user vectors based on text information) based on the tag information generated by the tag generation unit 1022 is also the same as the generation examples described above.
[0358] [Example of determining the ratio of image selection information and text information to be reflected based on user requests] Here, we show an example of generating a new user vector C for selecting an image, using a user vector (user vector A) generated based on image selection information and a user vector (user vector B) generated based on text information.
[0359] For example, it is possible to generate user vector C using a predetermined ratio (the reflection ratio of user vector A and user vector B). For instance, using the reflection ratio α of user vector A and user vector B, user vector C can be calculated using the following equation 7. Note that the calculation method using equation 7 is just one example, and user vector C may be calculated using other calculation methods. User vector C = (1-α) × User vector A + α × User vector B ...Equation 7
[0360] Furthermore, the vector generation unit 1023 generates user vector A and user vector B based on the tag information output from the tag generation unit 1022, and generates user vector C based on the reflection ratio α and user vectors A and B. The reflection ratio α may be set by the user or by the service provider in this embodiment.
[0361] Furthermore, the reflection ratio α may be set based on at least one of the text information. For example, it is possible to set the reflection ratio α using an AI model. For instance, it is possible to use an AI model that has learned from multiple text information sets to which predetermined tags associating the reflection ratio with the request text are assigned as training labels.
[0362] Furthermore, it is possible to use an AI model such as LLM. Examples of LLMs that can be used include ChatGPT, Bard, Llama2, Gemini, and Claude. Note that these are just examples, and other AI models may be used. For example, conditional information indicating that the system should output a reflection ratio α based on text information entered by the user can be input to the LLM as instruction information (prompt), and the output information in response to this can be the reflection ratio α.
[0363] Here, we show an example of a prompt that outputs the reflection ratio α using ChatGPT. For example, it is possible to input instruction information (prompt) to the LLM to determine the degree to which the content of the text information transmitted from the electronic device 1100 should be reflected, and output that percentage (0 to 1), and the output information in response to this is the reflection ratio α. In this case, instruction information (prompt) to output the basis for estimating the reflection ratio α may also be input to the LLM.
[0364] For example, as an example of how to determine the reflection ratio α (an example of the basis for estimating the reflection ratio α), if a user inputs, "Please erase the original recommendation results and recommend images of mugs with a bright atmosphere," it is clear that the user wants to ignore the original recommendation results, so it is possible to estimate that a reflection ratio of 1.0 is appropriate for the content of that sentence. Also, if a user inputs, for example, "Please recommend mugs with a slightly brighter atmosphere," it is clear that the user wants to maintain the original recommendation results to some extent while making some changes, so it is possible to estimate that a reflection ratio of around 0.3 is appropriate for the content of that sentence. Furthermore, if a user inputs, for example, "I like mugs with a bright atmosphere," it is difficult to determine whether or not to maintain the original recommendation results, so it is possible to estimate that a reflection ratio of around 0.5 is appropriate for the reflection ratio α.
[0365] [Example of image selection] Next, the extraction unit 1024 extracts images from the image information DB 1040 to be newly provided to the user based on the user vector C generated by the vector generation unit 1023. This extraction method is the same as the method for extracting new images when the new information re-suggestion button 1233 (see Figure 27) is selected, so a detailed explanation is omitted here. For example, the extraction process for new images is performed using a category (e.g., mug) held in the memory of the information processing device 1000.
[0366] Furthermore, the extraction unit 1024 outputs image information (for example, image information 1042) related to the extracted image to the information provision unit 1025. The information provision unit 1025 then transmits the image information to the electronic device 1100, which displays the image information.
[0367] The above example illustrates how to extract images to provide to a user based on the user vector C. However, based on the user vector C, providers can also be extracted from the provider information DB 1050 according to the user's preferences. In this case, provider information (e.g., profile information 1052, image information 1042) related to the extracted providers is transmitted to the electronic device 1100 and displayed on the electronic device 1100. For example, the provider information screen 1230 (see Figure 27) is displayed on the UI unit 1160.
[0368] Alternatively, a text information input field 1405 (see Figures 36 and 39) may be provided on the provider information screen 1230 (see Figure 27). The text information entered in the text information input field 1405 may be used to generate a user vector B that corresponds to the user's preferences. The image extraction process, provider extraction process, etc., may then be performed using the already generated user vector A (a user vector generated based on image selection information) and user vector B. In this case, each extracted piece of information is transmitted to the electronic device 1100 and displayed on the electronic device 1100.
[0369] [Example of selecting a new image] Figure 40 shows an example of the display of the selection screen 1440 shown on the UI unit 1160 of the electronic device 1100. The selection screen 1440 is displayed on the UI unit 1160 based on the control of the control unit 1150. The selection screen 1440 is a modified version of the selection screen 1240 shown in Figure 28, with a modified portion of the upper message; the rest is the same as the selection screen 1240. Therefore, the parts common to the selection screen 1240 are denoted by the same reference numerals. The image display area 1241 displays each image selected by the image selection process based on the user vector C described above, but here, for the sake of simplicity, an example is shown in which multiple images 1243 to 1250 similar to those on the selection screen 1240 are displayed.
[0370] In this way, it is possible to select and provide a new image to the user using a user vector C generated based on the text information entered by the user and the image selected by the user. Furthermore, a text information input field 1405 (see Figures 36 and 39) may be provided on the selection screen 1440 to allow for additional text information input. In this case, when text information is entered by the user, it is possible to provide the user with yet another new image based on that text information. Moreover, by performing each of these processes multiple times, it is possible to refine the recommendations provided to the user.
[0371] [Example of server operation] For an example of the operation of the information processing device 1000, the selection process shown in Figure 38 can be applied. However, in step S1302, it is determined whether or not only image selection information has been received. In step S1352, a user vector is generated based on the image selection information and text information (or text information only). That is, a user vector C for selecting a new image is generated using the user vector generated based on the image selection information (user vector A) and the user vector generated based on the text information (user vector B). In this case, it is possible to generate the user vector C using the equation 7 described above.
[0372] In this way, when a user sends information about their preferences (e.g., image selection information, text information) to the information processing device 1000, the information processing device 1000 generates a user vector (e.g., user vector C) based on each piece of information, and based on that user vector, it is possible to select and present images that match the user's preferences to the user. In this case, if there is a favorite image among the images presented to the user, that favorite image is selected and sent to the information processing device 1000. In this case, it is possible to add a new favorite image in addition to the favorite image that the user has already selected, thereby generating a user vector that better reflects the user's preferences. In this way, by sequentially adding the user's favorite images, it is possible to update the user vector to better reflect the user's preferences and improve the accuracy of image recommendation. On the other hand, if there is no favorite image among the images presented to the user, the user sends a request (e.g., text information) using other information, and the information processing device 1000 can reflect that request and recommend images again. Note that the image selection information also includes selection information when a favorite image is selected from among randomly displayed images.
[0373] The processes shown in Figures 29, 32, 35, and 38 are merely examples for realizing this embodiment, and the order of some of the processing steps may be rearranged, some of the processing steps may be omitted, or other processing steps may be added, as long as the embodiment is feasible. Furthermore, while the examples shown in these processes illustrate the control unit 1150 controlling the display state of the display screen shown on the electronic device 1100 based on information transmitted from the information processing device 1000, the embodiment is not limited to this. For example, the display state of the display screen shown on the electronic device 1100 may be controlled based on control from the information processing device 1000.
[0374] In this way, the user can request improvements to the suggested image from the information processing device 1000, and based on the content of the request, a request vector (user vector B) can be generated, and the user vector can be updated using this request vector. This updated user vector is user vector C. By updating the user vector in this way based on the content of the user's request, the user can communicate their requests more specifically and have a more direct influence on the image recommendation results. As a result, it is possible to improve user satisfaction and maintain the user's interest.
[0375] [Example of generating tag information using text information converted from an image] As shown in the first and second embodiments, it is possible to generate text information based on an image and extract multiple feature quantities corresponding to each of the multiple tags based on that text information. Therefore, in the third embodiment, we will show an example of applying the first and second embodiments. Note that some of the explanations below will be omitted for parts that are common to the first and second embodiments.
[0376] For example, if one or more images are selected in the image selection screen shown in Figures 26, 36, 37, 39, and 40, the tag generation unit 1022 generates text information based on the selected one or more images. For example, if one image is selected, it generates text information corresponding to that image. Also, for example, if multiple images are selected, it generates multiple pieces of text information corresponding to each of those images. Then, based on the generated text information, the tag generation unit 1022 extracts multiple feature quantities corresponding to each of the multiple tags for each tag. Note that even if tag information is associated with the selected image in the image information DB 1040, text information may be generated based on the selected image, and tag information may be generated based on that text information (for example, the process in S1303).
[0377] In this case, as in the first and second embodiments, multiple tags may be set based on target information specified by user operation (e.g., a category such as a mug), multiple tags may be set based on generated text information, or objects contained in the image (e.g., a mug) may be detected and multiple tags may be set based on the detected objects.
[0378] Furthermore, similar to the first and second embodiments, multiple text information may be generated based on a single image, and multiple feature quantities corresponding to each of the multiple tags may be extracted for each tag based on that multiple text information.
[0379] The vector generation unit 1023 then generates a user vector using multiple tags generated based on the text information. This user vector is user vector A because it was generated based on the image selection information. The method for generating a user vector using multiple tags is the same as the generation method described above. Alternatively, as in the example shown in Figures 36 to 40, a user vector C for selecting a new image may be generated using user vector A and a user vector (user vector B) generated based on text information (for example, the user's request). The extraction unit 1024 then extracts an image or provider using the user vector (user vector C) generated by the vector generation unit 1023.
[0380] In this way, by extracting feature quantities for each tag based on text information generated from the image, it is possible to extract new feature quantities that differ from those extracted directly from the image. In other words, it is possible to extract new feature quantities that take into account differences arising from differences between languages, etc. This makes it possible to appropriately extract image features, allowing users to more appropriately select their preferred images or providers.
[0381] [Example of setting the reflection ratio using the similarity of user vectors] Alternatively, for example, the similarity between a user vector (user vector A) generated based on image selection information and a user vector (user vector B) generated based on text information may be calculated, and user vector C may be generated based on that similarity. For example, it is possible to calculate the difference value (for each corresponding component) between each component constituting user vector A and each component constituting user vector B, and calculate this difference value as the similarity. Alternatively, for example, the cosine similarity between user vector A and user vector B can be calculated as the similarity. As mentioned above, it is preferable to normalize each vector when calculating cosine similarity. Furthermore, a high similarity between user vector A and user vector B means that user vector A and user vector B are close in the vector space.
[0382] Furthermore, it is possible to set the reflection ratio α based on the similarity between user vector A and user vector B. For example, if the similarity is high (for example, if the similarity is equal to or greater than the first criterion), it means that an image close to the user's request is being recommended. Therefore, it is possible to set the reflection ratio of the content of the user's request text to a value close to 0. On the other hand, if the similarity is low (for example, if the similarity is less than the second criterion (however, the first criterion > the second criterion)), it means that an image different from the user's request is being recommended. Therefore, it is possible to set the reflection ratio of the content of the user's request text to a value close to 1. Also, in the case of moderate similarity (for example, if the similarity is equal to or greater than the second criterion but less than the first criterion), it is possible to set the reflection ratio of the content of the user's request text within the range of 0 to 1 according to the similarity. In this way, it is possible to calculate the similarity between user vector B (first feature) and user vector A (second feature), and then calculate the weight (reflection ratio α) between user vector A and user vector B based on that similarity. In other words, it is possible to check to what extent the recommended images reflect the content of the user's request text, and then provide the user with new images based on the results of that check.
[0383] [Example of recommending images using vectors that are not similar to user vectors] The above examples illustrate how to provide a user with images similar to at least one of user vectors A-C. However, it is conceivable that a user's interests and preferences may differ from those initially assumed. For example, if a user who prefers modern designs is shown classic or retro mugs—the opposite of their preference—they may develop a strong interest in classic or retro mugs. Therefore, this section presents an example of presenting a user with an image of a vector with a different direction (e.g., in the opposite direction) from their user vector (at least one of user vectors A-C) as reference information. For example, it is possible to use a vector D with low similarity to the user vector (at least one of user vectors A-C) (e.g., a vector whose distance from the user vector (at least one of user vectors A-C) in vector space is greater than or equal to a certain threshold). As vector D, for example, it is possible to use a vector obtained by transforming the user vector (at least one of user vectors A-C) symmetrically with respect to an average vector (e.g., a vector that is the exact opposite of the user vector). Alternatively, a correlation database (CDB) could be prepared to store information about the correlation between each tag (for example, information indicating the relationship between multiple tags, such as a high value for Classic and a low value for Modern), and vector D could be generated using this correlation database. By using this correlation database, it is possible to calculate, for example, the exact opposite of the user vector. The contents of the correlation database may be adjusted as appropriate according to the preferences of the administrator or user. Furthermore, these examples are just examples, and vector D may be obtained by other transformation methods. For example, in the vector space, the vector furthest from the user vector (or a vector within a predetermined distance from the furthest vector relative to the user vector, or a vector that is more than a predetermined distance from the user vector, etc.) could be set as vector D. In these cases, an image with a different direction from the user vector (at least one of user vectors A to C) (for example, the exact opposite image) can be selected based on vector D.In other words, it is possible to present the user with images that are in a direction different from the user vector (at least one of user vectors A to C) (for example, the exact opposite image). In this case, it is possible to display one or more images (first image) selected based on the user vector (at least one of user vectors A to C) and one or more images (second image) selected as images that are in a direction different from the user vector (at least one of user vectors A to C) (for example, the exact opposite image), in a way that the user can compare them. For example, it is possible to display a first image display area that displays one or more first images and a second image display area that displays one or more second images side by side, either vertically or horizontally. Furthermore, it may be explicitly stated in the first image display area or its vicinity that the images are to the user's liking, and in the second image display area or its vicinity that the images are in a direction different from the user's liking (for example, the exact opposite image). This allows the user to see unexpected images and is expected to make new discoveries.
[0384] Alternatively, the decision to provide the second image to the user may be made based on the similarity between user vector A and user vector B. For example, if the similarity is high (for example, if the similarity is equal to or greater than the first criterion), it means that an image determined to be close to the user's preference has been provided, but the user has entered a request. Therefore, it is assumed that the user's preference determined by the information processing device 1000 differs from the actual user's preference (request). In such cases, it is possible to provide the user with both the first image (the image determined to be the user's preference by the information processing device 1000) and the second image (the opposite image) to observe the user's reaction (the user's actual preference). On the other hand, if the similarity is low (for example, if the similarity is less than the second criterion), the information processing device 1000 can select a new image that the user prefers (the first image) and provide only the first image to the user to observe the user's reaction.
[0385] Furthermore, if a second image is displayed, it is anticipated that the second image may be selected as the user's preferred image. In this case, the weight β3 of vector D can be increased to generate a new user vector C (for example, user vector C = β1 × user vector A + β2 × user vector B + β3 × user vector D ... Equation 8, where β1~β3≧0, β1+β2+β3=1). The relationship between β1 and β3 can be set, for example, according to the ratio of the number of first images selected by the user to the number of second images selected by the user after the first and second images are displayed. For example, if the number of first images selected by the user is 3 and the number of second images selected is 2, then β1:β3 can be set to 3:2. Also, for example, if no new request text is entered by the user after the first and second images are displayed, β2 is set to 0. On the other hand, if new request text is entered by the user, β1~β3 can be set based on that new request text. In this case, as described above, the reflection ratios β1 to β3 can be set using an AI model (e.g., LLM). Note that these weights are just examples, and weights obtained by other calculations may also be used. That is, the information processing device 1000 can include a search unit (e.g., tag generation unit 1022, vector generation unit 1023, extraction unit 1024) that searches for a first image based on the user's preferred orientation and a second image based on an orientation different from the said orientation (e.g., the opposite direction, or an orientation that differs from that orientation by a predetermined value or more (e.g., an angle that differs by a predetermined value or more)), and a control unit (e.g., information provision unit 1025) that displays the first image and the second image on a display unit (e.g., UI unit 1160 of electronic device 1100) in a display mode that allows comparison of the two images. Furthermore, the first and second images described above may be searched for and displayed in the display unit within the range of categories set on the initial search screen (for example, search screen 1200 (see Figure 25)) (for example, categories narrowed down in advance based on the entered text information (for example, "I want a stylish mug") (for example, "mug")), or the first image found within the range of categories set on the initial search screen and the second image found outside the range of that category may be displayed in the display unit. The range outside the category can be set, for example, based on each category stored in a pre-configured category DB. For example, a combination of opposing categories or completely different categories can be pre-configured, and the range outside the category can be set based on this combination. For example, the category "mug" and the category "clock" can be stored in the category DB as a combination of different categories, and based on this category DB, it is possible to search for the second image within the range of the category "clock" which is outside the range of the category "mug". In this case, from within the category "Clock," it is possible to select as the second image (one or more images) an image with the same orientation as the user vector (at least one of user vectors A to C) described above, or an image with a different orientation (for example, an image that is the exact opposite).
[0386] [Example of the effects of the third embodiment] Thus, in this third embodiment, by utilizing techniques such as statistical machine learning and analyzing images selected by the user (for example, product images), it is possible to provide or recommend images that the user prefers in terms of design. For example, by designing tags related to the design taste of product images and training a machine learning model for each category, tags can be automatically assigned. Then, by having the user select their favorite images and using the tags assigned to them, it is possible to quantify their preferences (user vector). By using this user vector and calculating the similarity between the image and the provider, it is possible to recommend new images that each person is likely to like. In other words, based on the user's favorite images selected from the perspective of design taste (for example, product images), it is possible to recommend images with similar design tastes.
[0387] Furthermore, by utilizing LLM, it is possible to re-recommend recommended images based on user feedback. For example, a user can submit a text-based request to LLM for improvement on a recommended image (e.g., "I prefer a more modern design"), and LLM can generate a request vector based on the request and update the user vector using that vector. This allows users to communicate their requests more specifically and to have a more direct impact on the recommendation results. As a result, user satisfaction can be improved and interest can be sustained. In this way, by introducing a text-based feedback mechanism using LLM into the system, it becomes easy for users to provide feedback on products recommended to them.
[0388] In this way, when a user selects an image they like, it is possible to recommend images similar in style to the selected image, as well as providers that match that style. In other words, it is possible to recommend images that include products or services with a design style similar to the design style of the image the user selected (the design style of the product or service). This makes it possible, for example, on an e-commerce site to recommend products by understanding the abstract ideal image or preference that the user has regarding the design style of a product. Furthermore, if the user is not satisfied with the recommended product, they can easily provide feedback to the system and generate new recommendations.
[0389] This makes it easier for users to find images that match their design preferences (for example, images of products that suit their taste), thus increasing the number of transactions conducted through e-commerce sites. This is considered particularly effective for the sale of products that are selected based on design taste (for example, houses, clothing, furniture, miscellaneous goods, tableware, houseplants, paintings, etc.).
[0390] Thus, in this third embodiment, it is possible to recommend products or services by taking into account the abstract ideal image and preferences that the user has regarding the design taste of the product or service. Furthermore, if the user is not satisfied with the recommended product or service, the user can provide feedback to the system to improve the recommendation. In other words, it is possible to provide information appropriately tailored to the user's preferences.
[0391] [Example of an information processing system configuration] The above examples illustrate how tag setting, text information generation, feature extraction, tag generation, vector generation, and extraction processes are performed on information processing devices 10, 50, 500, 560, 800, 900, 1000, etc. However, all or part of these processes may be performed on other devices. In this case, the information processing system is comprised of devices that perform parts of these processes. For example, at least part of each process can be performed using various information processing devices such as servers, user-accessible devices (e.g., smartphones, tablet devices, personal computers), and servers that can be connected via a predetermined network such as the Internet, as well as various electronic devices.
[0392] Furthermore, some (or all) of the information processing systems capable of performing the functions of information processing devices 10, 50, 500, 560, 800, 900, 1000, etc., may be provided by applications that can be provided via a predetermined network such as the Internet. This application may be, for example, SaaS (Software as a Service).
[0393] [Example configuration and effects of this embodiment] As described above, the configurations of the information processing devices 10, 50, 500, 560, 800, 900, and 1000 can be combined in any way as needed, in addition to the combinations described above. Therefore, examples considering these combinations will be shown below.
[0394] Information processing system IS1 is an information processing system that includes an electronic device 1100 used by a user and an information processing device 1000 capable of providing image information to the electronic device 1100. Furthermore, information processing system IS1 is an information processing device capable of providing the electronic device 1100 with recommended images (for example, images 1402-1404 (see Figure 39)) that have been retrieved using selected images chosen according to the user's preferences. The information processing system IS1 includes an image information DB 1040 (an example of a database) that stores images and their features in association with each of multiple images, an information acquisition unit 1021 (an example of an acquisition unit) that acquires text information indicating requests for recommended images (for example, characters in the text information input field 1405 (see Figure 39)) (or an information acquisition unit 1021 that acquires selected images selected according to the user's preferences (for example, images 1402 and 1403 (see Figure 39)) and text information indicating requests for those selected images (for example, characters in the text information input field 1405 (see Figure 39))), and generates first features related to the user's preferences based on the text information, and based on the selected images The system includes a tag generation unit 1022 that generates features (second features) of the selected image related to the items included in the selected image (e.g., a mug), a vector generation unit 1023 (an example of a feature generation unit; however, the weights of the first and second features (e.g., reflection ratio α) may be generated based on text information), and an extraction unit 1024 (an example of a selection unit) that selects a new image from among the images stored in the image information DB 1040 that matches the user's preferences based on the comparison result of comparing feature information indicating the user's preferences (e.g., user vector C) generated based on the first and second features (or using weights). The system also includes a provider information DB 1050 (an example of a database) that stores features related to products or services provided by providers in association with multiple providers, and the extraction unit 1024 may select a provider that matches the user's preferences from among the providers stored in the provider information DB 1050 based on the comparison result of comparing feature information indicating the user's preferences (e.g., user vector C) with features related to products or services.The second feature is tag information (or the corresponding user vector A) generated based on the image information, and the first feature is tag information (or the corresponding user vector B) generated based on the text information. Furthermore, user vector C is calculated using Equation 7 based on user vector A, user vector B, and the reflection ratio α. In other words, the vector generation unit 1023 generates feature information (e.g., user vector C) indicating the user's preferences based on text information indicating the request for the recommended image (e.g., the first feature) and the selected image selected according to the user's preferences (e.g., the second feature). In this case, it is possible to generate feature information (e.g., user vector C) indicating the user's preferences using the respective weights (e.g., reflection ratio α) of the first and second features generated based on the text information. That is, the information processing device 1000 can select a new image according to the user's preferences based on text information indicating the request for the recommended image (e.g., the first feature) and the selected image selected according to the user's preferences (e.g., the second feature). Furthermore, the second feature may be that text information is generated based on the selected image, and tag information is generated based on that text information. Also, the information processing method according to this embodiment is an information processing method that includes each of these processes. Furthermore, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each of the functions that the information processing device 1000 can execute.
[0395] This configuration allows for the selection of new images tailored to the user's preferences using a user vector C generated based on text information entered by the user and images selected by the user. Therefore, it becomes possible to suggest appropriate images that match the user's preferences. This makes it possible, for example, to accurately understand the user's abstract ideal image and preferences regarding the design style of products and services, and to recommend such products and services to the user. Furthermore, it enables intuitive and easy-to-understand support for various user selections.
[0396] The information processing system IS1 is an information processing system that includes an electronic device 1100 used by a user and an information processing device 1000 capable of providing image information to the electronic device 1100. The information processing system IS1 includes an image information DB 1040 (an example of a database) that stores images and their features in association with each of multiple images; an information acquisition unit 1021 (an example of an acquisition unit) that acquires selected images from the electronic device 1100 according to the user's preferences; a tag generation unit 1022 (an example of a text information generation unit) that generates text information based on the selected images; a tag generation unit 1022 (an example of a feature generation unit) that generates features corresponding to predetermined items (one or more items) based on the text information (for example, extracting multiple feature quantities (examples of features) corresponding to each of multiple tags (examples of items) for each item) to be used as features of the selected images; and an extraction unit 1024 (an example of a selection unit) that newly selects an image according to the user's preferences from among the images stored in the image information DB 1040 based on the comparison result of comparing the features of the selected image (for example, user vector A or C) with the features of the image (for example, image vector). Furthermore, the information processing method according to this embodiment is an information processing method that includes each of these processes. Furthermore, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to implement each of the functions that the information processing device 1000 can perform.
[0397] This configuration allows for the extraction of feature quantities for each tag based on text information generated from an image. For example, a general-purpose AI model can be used in the feature generation unit. In this case, a large amount of training data is not required to extract image features. That is, it becomes possible to appropriately extract image features without preparing a large amount of training data. Furthermore, because feature quantities for each tag are extracted based on text information generated from the image, it is possible to extract new features that differ from those extracted directly from the image. That is, it is possible to extract new features that take into account differences arising from differences between languages, etc. In this way, it becomes possible to appropriately extract image features. Moreover, using the extracted image features, it is possible to select a new image from among multiple images that matches the user's preferences. Therefore, it becomes possible to suggest appropriate images that match the user's preferences. This makes it possible, for example, to appropriately understand the abstract ideal image and preferences that the user has regarding the design taste of products and services, and to recommend such products and services to the user. It also makes it possible to support various user choices in an intuitive and easy-to-understand manner.
[0398] Information processing system IS1 is an information processing system that includes an electronic device 1100 used by a user and an information processing device 1000 capable of providing image information to the electronic device 1100. Information processing system IS1 includes a provider information DB 1050 (an example of a database) that stores characteristics of goods or services provided by a provider and the provider associated with each of several providers, an information acquisition unit 1021 (an example of an acquisition unit) that acquires selected images from the electronic device 1100 according to the user's preferences, a tag generation unit 1022 and a vector generation unit 1023 (an example of a feature generation unit) that generate features of the selected image related to what is included in the selected image (e.g., a mug) based on the selected image, and an extraction unit 1024 (an example of a selection unit) that selects a provider according to the user's preferences from among several providers stored in the provider information DB 1050 based on a comparison result obtained by comparing the features of the selected image with the features of goods or services. Furthermore, the information processing method according to this embodiment is an information processing method that includes each of these processes. Furthermore, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to implement each of the functions that the information processing device 1000 can perform.
[0399] This configuration allows the system to acquire selection images from the electronic device 1100 according to the user's preferences, and then, based on a comparison of the characteristics of the acquired images (e.g., mugs) with the characteristics of the products or services in the provider information DB 1050, select a provider from among multiple providers that best suits the user's preferences. This makes it possible to suggest appropriate providers that match the user's preferences. For example, this makes it possible to appropriately understand the user's abstract ideal image or preferences regarding the design taste of products and services, and recommend such products and services to the user. Furthermore, it makes it possible to support the user's various selections in an intuitive and easy-to-understand manner.
[0400] The information acquisition unit 1021 (an example of an acquisition unit) acquires multiple selection images from the electronic device 1100 according to the user's preferences. The information processing system IS1 further includes a vector generation unit 1023 (an example of a feature information generation unit) that generates a user vector (an example of feature information) indicating the user's preferences based on the characteristics of each of the multiple selection images. The extraction unit 1024 (an example of a selection unit) selects a new image according to the user's preferences based on the comparison result obtained by comparing the user vector with the image features (e.g., image vectors) of the image information DB 1040 (an example of a database).
[0401] This configuration allows for the selection of a new image from among multiple images that matches the user's preferences, based on a comparison result obtained by comparing user vectors, which are generated based on the characteristics of each of the user's preferred images, with the characteristics of the images in the image information DB1040. Therefore, it becomes possible to suggest a more appropriate image that matches the user's preferences.
[0402] The information processing system IS1 further includes an information providing unit 1025 that transmits display information to the electronic device 1100 for displaying the newly selected image by the extraction unit 1024 (an example of a selection unit). The electronic device 1100 includes a control unit 1150 that performs display control to display the newly selected image on the UI unit 1160 (an example of a display unit) based on the display information.
[0403] With this configuration, based on the user's preferred images, the user can easily visually confirm the newly selected image from among multiple images that matches their preferences on the UI unit 1160 of the electronic device 1100.
[0404] The information processing system IS1 further includes an information providing unit 1025 that transmits display information for displaying images stored in the image information DB 1040 (an example of a database) to the electronic device 1100. The information acquisition unit 1021 (an example of an acquisition unit) acquires from the electronic device 1100, based on the display information, images selected by the user based on user operation, as selected images according to the user's preference.
[0405] This configuration makes it possible to display images stored in the image information DB 1040 on the electronic device 1100. Therefore, users can easily visually confirm the images displayed on the electronic device 1100 and easily select their preferred image from among them.
[0406] The electronic device 1100 includes a control unit 1150 that performs display control to display an image stored in the image information DB 1040 (an example of a database) on the UI unit 1160 (an example of a display unit) based on display information, and transmission control to transmit edited image information related to the edited image, which is an edited image based on user operation, to the information processing device 1000. In addition, the information acquisition unit 1021 (an example of an acquisition unit) acquires an image corresponding to the edited image information transmitted from the electronic device 1100 as a selected image.
[0407] This configuration allows users to edit images displayed on the electronic device 1100 according to their preferences, thereby enabling a more appropriate expression of user preferences based on user-initiated editing. Therefore, it becomes possible to suggest appropriate images and other elements based on the user's editing results, tailored to their preferences.
[0408] When the electronic device 1100 sends a re-suggestion request requesting the re-suggestion of new images other than the newly selected image, the information provision unit 1025 uses the features generated by the tag generation unit 1022 and the vector generation unit 1023 (an example of a feature generation unit) to send display information to the electronic device 1100 for displaying the newly selected image from among the images stored in the image information DB 1040 (an example of a database) that matches the user's preferences.
[0409] This configuration allows the user to use the electronic device 1100 to submit a re-suggestion request, requesting new images other than those initially suggested. This makes it easy for the user to request re-suggestions of images that suit their preferences, even if an image unsuitable for their taste is suggested. In this case, the image used for the new selection can be chosen based on the characteristics of the image previously selected by the user, thus providing the user with an image closer to their preferences. Therefore, it becomes possible to suggest more appropriate images that align with the user's preferences.
[0410] The electronic device 1100 includes an image information DB 1080 (an example of an image information database) that stores images acquired by the image acquisition unit 1130, a control unit 1150 that performs display control to display images stored in the image information DB 1080 on the UI unit 1160 (an example of a display unit), and transmission control to transmit selected image information to the information processing device 1000 regarding an image selected from among the images displayed on the UI unit 1160 based on user operation. In addition, the information acquisition unit 1021 (an example of an acquisition unit) acquires an image corresponding to the selected image information transmitted from the electronic device 1100 as a selected image.
[0411] This configuration allows users to select images according to their preferences from those acquired using the electronic device 1100 (or other external devices), thereby more appropriately expressing user preferences based on the user's proactive image collection. Therefore, it becomes possible to suggest appropriate images and other content based on the user's collection results, according to the user's preferences.
[0412] The tag generation unit 1022 (an example of a feature generation unit) may generate at least one of the first feature, the second feature, and a weight using an AI model that extracts features related to what is included in the target image.
[0413] This configuration makes it possible to appropriately generate first features, second features, weights, etc., for any image transmitted from the electronic device 1100 according to the user's preferences. Therefore, it becomes possible to suggest appropriate images, etc., according to the user's preferences.
[0414] The information processing device 1000 is capable of providing recommended images, retrieved using selected images chosen according to the user's preferences, to the electronic device 1100 used by the user, and is capable of handling an image information DB 1040 (an example of a database) that stores images and their characteristics in association with each of multiple images. The information processing device 1000 includes an information acquisition unit 1021 (an example of an acquisition unit) that acquires text information indicating requests for recommended images; a tag generation unit 1022 and a vector generation unit 1023 (an example of a feature generation unit) that generate a first feature relating to the user's preferences based on the text information, generate a second feature relating to what is included in the selected image (e.g., a mug) based on the selected image, and generate the respective weights (e.g., reflection ratio α) of the first and second features based on the text information; and an extraction unit 1024 (an example of a selection unit) that newly selects an image from among the images stored in the image information DB 1040 that matches the user's preferences based on the comparison result obtained by comparing the feature information indicating the user's preferences (e.g., user vector C) generated based on the weights, the first feature, and the second feature with the features of the image (e.g., an image vector). The first feature is tag information (or the corresponding user vector A) generated based on the image information, and the second feature is tag information (or the corresponding user vector B) generated based on the text information. Furthermore, based on user vector A, user vector B, and reflection ratio α, user vector C is calculated by equation 7. The information processing method according to this embodiment is an information processing method that causes a computer to execute each of these processes. The program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each of the functions that the information processing device 1000 can execute.
[0415] [Example configurations and effects of the first and second embodiments] Conventionally, there are technologies that perform various processes using various information related to images. For example, a technology has been proposed that uses a trained model to extract multiple features corresponding to each of multiple images (for example, Japanese Patent Publication No. 2023-000313).
[0416] Conventional techniques allow for the extraction of image features using pre-trained models. For example, generating a pre-trained model to extract image features in a certain direction requires a large amount of training data where features have been associated through human work. However, preparing a large amount of training data requires many personnel with specialized knowledge, and generally, it is often difficult to obtain such a large number of personnel. Furthermore, if a large amount of training data cannot be prepared, it is expected that it will be difficult to properly extract image features.
[0417] Therefore, in this embodiment, by using the following configuration example, it becomes possible to appropriately extract image features.
[0418] Information processing devices 50, 500, 560, 800, and 900 are examples of information processing systems that extract feature quantities (examples of features) related to tags (examples of items) for items contained in an image (e.g., a dress). For example, information processing devices 800 and 900 include a tag setting unit 12 (an example of a setting unit) that sets multiple tags based on target information (e.g., a category such as clothing) specified by user operation, an image acquisition unit 51 and a communication unit 910 that acquire an image, a text information generation unit 520 (corresponding to the text information generation unit 520 shown in Figures 9 and 15) (an example of a generation unit) that generates multiple text information related to items contained in the image (e.g., a dress), a feature extraction unit 530 (corresponding to the feature extraction unit 530 shown in Figures 9 and 15) (an example of an extraction unit) that extracts multiple feature quantities corresponding to each of the multiple tags from the multiple text information as a criterion for extracting feature quantities from the target information (e.g., from the perspective of a clothing expert) for each tag, and a DB control unit 810 (an example of a control unit) that outputs the image and the multiple feature quantities extracted for each tag in association. The information processing devices 800 and 900 described herein may consist of one device or multiple devices. Alternatively, instead of the information processing devices 800 and 900, an information processing system consisting of multiple devices capable of performing the processes realized by the information processing devices 800 and 900 may be used.
[0419] This configuration allows for the extraction of feature quantities for multiple tags based on text information generated from an image, tag by tag. For example, a general-purpose AI model can be used in both the text information generation unit and the feature extraction unit. In this case, a large amount of training data is not required to extract image features. That is, it becomes possible to appropriately extract image features without preparing a large amount of training data. Furthermore, because feature quantities for multiple tags are extracted tag by tag based on text information generated from an image, it is possible to extract new features that differ from those extracted directly from the image. That is, it is possible to extract new features that take into account differences arising from differences between languages, etc. In addition, since it is possible to extract multiple features from text information based on target information specified by the user (for example, a category such as clothing) (for example, from the perspective of a clothing expert), it becomes possible to extract appropriate features according to the user's preferences. That is, it becomes possible to appropriately extract image features.
[0420] The information processing devices 50, 560, 800, and 900 include an image acquisition unit 51 and a communication unit 910 (an example of an acquisition unit) that acquire an image 41; text information generation units 52 and 520 (an example of generation units) that generate text information 42 based on the image; feature extraction units 53 and 530 (an example of extraction units) that extract features corresponding to predetermined items (one or more items) based on the text information (for example, extracting multiple feature quantities (an example of features) corresponding to each of multiple tags (an example of items) for each item); and a recording control unit 54 and a DB control unit 810 (an example of a control unit) that associate the image with the multiple feature quantities extracted for each item. When one tag is set by the tag setting unit 12 (an example of a setting unit), the feature extraction units 53 and 530 extract one feature quantity corresponding to one tag (an example of a predetermined item) based on the text information. Furthermore, the information processing method according to this embodiment is an information processing method that includes each of these processes. Furthermore, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to implement each of the functions that each information processing device can perform.
[0421] This configuration allows for the extraction of feature quantities for multiple tags based on text information generated from an image, tag by tag. For example, a general-purpose AI model can be used in both the text information generation unit and the feature extraction unit. In this case, a large amount of training data is not required to extract image features. In other words, it becomes possible to appropriately extract image features without preparing a large amount of training data. Furthermore, because feature quantities for multiple tags are extracted tag by tag based on text information generated from an image, it is possible to extract new features that differ from those extracted directly from the image. That is, it is possible to extract new features that take into account differences arising from differences between languages, etc. In this way, it becomes possible to appropriately extract image features.
[0422] The information processing devices 10 and 800 further include a tag setting unit 12 (an example of a setting unit) that sets multiple tags (an example of items) based on target information (for example, a category such as clothing) specified by user operation, and the feature extraction unit 53 may extract multiple feature quantities for each tag from the text information 42 as a criterion for extracting feature quantities (an example of features) from the target information (for example, from the perspective of a clothing expert).
[0423] With this configuration, it is possible to extract multiple features from text information 42 based on target information specified by the user (for example, a category such as clothing) (for example, from the perspective of a clothing expert), thus enabling the extraction of appropriate features according to the user's preferences.
[0424] The information processing device 800 may further include a tag setting unit 12 (an example of a setting unit) that sets multiple tags (an example of an item) based on text information 42.
[0425] This configuration makes it possible to generate multiple tags based on text information 42 generated from an image 41 specified by the user, allowing for the setting of appropriate tags according to the user's preferences and enabling the extraction of feature quantities corresponding to those tags.
[0426] The information processing device 800 may further include a tag setting unit 12 (an example of an object detection unit) that detects an object (for example, a dress) included in the image 41, and a tag setting unit 12 (an example of a setting unit) that sets a plurality of tags (an example of an item) based on the detected object.
[0427] This configuration allows for the generation of multiple tags based on an object (e.g., a dress) contained in an image 41 specified by the user. This enables the setting of appropriate tags according to the user's preferences and the extraction of feature quantities corresponding to those tags.
[0428] Furthermore, the text information generation unit 520 (an example of a generation unit) may generate multiple text information based on the image 41, and the feature extraction unit 530 (an example of an extraction unit) may extract multiple feature quantities (an example of features) for each tag (an example of an item) based on the multiple text information.
[0429] For example, using a single text information generation model (e.g., img2txt) results in a high degree of dependence on the characteristics of that model. However, by using multiple text information generation models (e.g., img2txt), it is possible to mitigate this dependence. Furthermore, using multiple models makes it possible to stabilize the behavior.
[0430] The text information generation unit 520 (an example of a generation unit) may generate multiple text information items through multiple generation processes using multiple algorithms.
[0431] This configuration makes it possible to generate multiple pieces of text information using multiple algorithms, and allows for stable system behavior regardless of the characteristics of each algorithm.
[0432] The feature extraction unit 530 (an example of an extraction unit) may extract multiple features for each tag (an example of an item) based on text information obtained by combining multiple pieces of text information.
[0433] This configuration makes it possible to extract multiple features for each tag based on text information that combines multiple pieces of text information, and it is possible to stabilize the system's behavior regardless of the features of each piece of text information.
[0434] The feature extraction unit 530 (an example of an extraction unit) may extract multiple feature quantities (an example of features) for each of the multiple text information items, for each tag (an example of an item). The information processing device 560 then processes the multiple feature quantities extracted for each of the text information items and the weight α i The system may further include a weighted average calculation unit 550 that calculates multiple feature quantities for each set of multiple tags by performing a predetermined first calculation process (for example, a weighted average calculation process) using the above.
[0435] According to this configuration, weight α i By using a weighted average calculation process, it is possible to obtain multiple features for each of the multiple tags, thereby improving the accuracy of feature calculations.
[0436] The text information generation unit 520 may generate multiple text information TX1 to TX4 based on one or a predetermined number of image information TD2s to which target information (e.g., categories such as clothing) related to multiple tags (an example of items) is set, and tag information TD3 (an example of feature information) indicating multiple feature quantities (an example of features) related to each of the multiple tags contained in the image information TD2 (e.g., dresses) is attached. Furthermore, the feature extraction unit 530 (an example of an extraction unit) may use the target information (e.g., categories such as clothing) as a criterion for extracting feature quantities (e.g., from the perspective of a clothing expert) and extract multiple feature quantities corresponding to each of the multiple tags from the multiple text information TX1 to TX4 for each tag. The information processing device 500 compares the multiple tag-specific feature quantities contained in the tag information TD3 attached to the image information TD2 with the multiple tag-specific feature quantities extracted by the feature extraction unit 530 for each tag, and based on a predetermined second calculation process (a calculation process that repeats (1) and (2) above M times) using the comparison result, it generates weights α related to the multiple text information TX1 to TX4. iA weight calculation unit 540 that performs calculations may be further provided.
[0437] According to this configuration, one or a small amount of tagged data TD1 is used to calculate the weights α used in the weighted average calculation. i It is possible to perform calculations appropriately. This makes it possible to improve the accuracy of feature calculations.
[0438] The information processing device 500 includes a text information generation unit 520 (an example of a generation unit) that generates multiple text information TX1 to TX4 based on one or a predetermined number of image information TD2s to which multiple tags (an example of items) related to target information (e.g., categories such as clothing) are set and tag information TD3 (an example of feature information) indicating multiple feature quantities (an example of features) for each of the multiple tags included in the image information TD2 (e.g., dresses), is attached; a feature extraction unit 530 (an example of an extraction unit) that extracts multiple feature quantities corresponding to each of the multiple tags from the multiple text information TX1 to TX4 for each tag as a criterion for extracting feature quantities related to the target information (e.g., categories such as clothing) related to the image information TD2 (e.g., categories such as clothing) (e.g., from the perspective of a clothing expert); and a weight α for the multiple text information TX1 to TX4 based on a predetermined calculation process (a calculation process that repeats (1) and (2) above M times) using the comparison result. i It comprises a weight calculation unit 540 that performs calculations. Furthermore, the information processing method according to this embodiment is an information processing method that includes each of these processes. Furthermore, the program according to this embodiment is a program that causes a computer to execute each of these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each of the functions that each information processing device can execute.
[0439] According to this configuration, one or a small amount of tagged data TD1 is used to calculate the weights α used in the weighted average calculation. iIt is possible to appropriately calculate α. In other words, it is possible to appropriately extract image features without preparing a large amount of training data. i Because it is possible to determine these features appropriately, the accuracy of feature extraction can be improved. In other words, it becomes possible to extract image features appropriately.
[0440] The processing steps shown in this embodiment are merely examples of how to implement this embodiment. The order of some of the processing steps may be changed, some of the processing steps may be omitted, or other processing steps may be added, as long as the embodiment is feasible.
[0441] Furthermore, each process shown in this embodiment is executed based on a program that causes a computer to execute each processing procedure. For this reason, this embodiment can also be understood as an embodiment of a program that realizes the function of executing each of these processes, and a recording medium that stores that program. For example, an update process to add a new function to an information processing device can cause that program to be stored in the storage device of the information processing device. This makes it possible to have the updated information processing device perform each of the processes shown in this embodiment.
[0442] Although embodiments of the present invention have been described above, these embodiments only represent a part of the application examples of the present invention, and are not intended to limit the technical scope of the present invention to the specific configurations of the above embodiments. For example, it is also possible to adopt the following configurations (configuration examples 1-13). [Configuration Example 1] An information processing system for extracting features related to items contained in an image, A setting unit that sets multiple items based on target information specified by user operation, The acquisition unit acquires the aforementioned image, A generation unit that generates multiple pieces of text information relating to the image, An extraction unit that uses the aforementioned target information as a criterion for extracting the aforementioned features, extracts the aforementioned features corresponding to each of the aforementioned items from the aforementioned plurality of text information, item by item, A control unit that associates the aforementioned image with the plurality of features extracted for each of the aforementioned items and outputs the result. An information processing system equipped with the following features. [Configuration Example 2] Image acquisition unit, A generation unit that generates text information based on the aforementioned image, An extraction unit that extracts features corresponding to predetermined items based on the aforementioned text information, A control unit that associates the image with the extracted features. An information processing device equipped with the following features. [Configuration Example 3] The aforementioned specified items consist of multiple items, The system further includes a setting unit that sets the multiple items based on target information specified by user operation, The extraction unit uses the target information as a criterion for extracting the features, and extracts the multiple features from the text information item by item. The information processing device described in Configuration Example 2. [Configuration Example 4] The aforementioned specified items consist of multiple items, The system further includes a setting unit that sets the multiple items based on the text information. The information processing device described in Configuration Example 2. [Configuration Example 5] The aforementioned specified items consist of multiple items, An object detection unit for detecting objects included in the aforementioned image, The system further comprises a setting unit that sets the plurality of items based on the detected object. The information processing device described in Configuration Example 2. [Configuration Example 6] The aforementioned specified items consist of multiple items, The generation unit generates multiple pieces of text information based on the image, The extraction unit extracts the multiple features item by item based on the multiple text information. An information processing device as described in any of configuration examples 2 to 5. [Configuration Example 7] The generation unit generates the multiple text information through multiple generation processes using multiple algorithms. The information processing device described in Configuration Example 6. [Configuration Example 8] The extraction unit extracts the multiple features item by item based on the text information obtained by combining the multiple text information. The information processing device described in Configuration Example 6. [Configuration Example 9] The extraction unit extracts the multiple features item by item from each of the multiple pieces of text information, The system further comprises a calculation unit that calculates the multiple features for each set of multiple items by performing a predetermined first calculation process using the multiple features and weights extracted for each piece of text information. The information processing device described in Configuration Example 6. [Configuration Example 10] The generation unit generates multiple text pieces based on one or a predetermined number of target images, each of which is assigned feature information indicating multiple features related to the multiple items and which are included in the target image. The extraction unit uses the target information as a criterion for extracting the features, and extracts a plurality of features corresponding to each of the plurality of items from the plurality of text information, item by item. The system further comprises a weight calculation unit that compares the features of each of the multiple items assigned to the target image with the extracted features of each of the multiple items for each item, and calculates the weights of the multiple text information based on a predetermined second calculation process using the comparison results. The information processing device described in Configuration Example 9. [Configuration Example 11] A generation unit that generates multiple text information based on one or a predetermined number of target images, each of which is assigned feature information indicating multiple features related to each of the multiple items included in the target image, and which is configured with multiple items related to the target information. An extraction unit that uses the aforementioned target information as a criterion for extracting the aforementioned features, extracts multiple features corresponding to each of the multiple items from the multiple text information for each item, A weight calculation unit compares the characteristics of each of the multiple items assigned to the target image with the extracted characteristics of each of the multiple items for each item, and calculates weights for the multiple text information based on a predetermined calculation process using the comparison results. An information processing device equipped with the following features. [Configuration Example 12] A generation process that generates text information based on an image, Based on the aforementioned text information, an extraction process is performed to extract features corresponding to predetermined items, Control process that associates the aforementioned image with the extracted features. Information processing methods including [Configuration Example 13] Generating text information based on an image, Based on the aforementioned text information, extract features corresponding to predetermined items, To associate the aforementioned image with the extracted features and A program that causes a computer to execute something. [Explanation of symbols]
[0443] 10, 50, 500, 560, 800, 900, 920 Information processing device, 11 Information acquisition unit, 12 Tag setting unit, 13, 54 Recording control unit, 14, 55 Storage unit, 51 Image acquisition unit, 52, 520 Text information generation unit, 53, 530 Feature extraction unit, 100 Tag DB, 200 Image DB, 510 Acquisition unit, 540 Weight calculation unit, 550 Weighted average calculation unit, 600 Weight DB, 810 DB control unit, 820 Output unit, 821, 931 Display unit, 910 Communication unit, NW1 Network, 930 Electronic device, 1000 Information processing device, 1010 Communication unit, 1020 Control unit, 1021 Information acquisition unit, 1022 Tag generation unit, 1023 Vector generation unit, 1024 Extraction unit, 1025 Information provision unit, 1030 Storage unit, 1040 Image information DB, 1050 Provider information DB, 1060 Tag information DB, 1100 Electronic equipment, 1110 Communication unit, 1120 Location information acquisition unit, 1130 Image acquisition unit, 1140 Sound acquisition unit, 1150 Control unit, 1160 UI unit, 1161 Reception unit, 1162 Output unit, 1170 Storage unit, IS1... Information processing system
Claims
1. An information processing system comprising an electronic device used by a user and an information processing device capable of providing the electronic device with recommended images retrieved using selected images chosen according to the user's preferences, A database that stores images and their characteristics in association with each set of images, An acquisition unit that acquires text information indicating requests for the aforementioned recommended images, A feature generation unit generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection unit selects a new image from the images stored in the database that matches the user's preferences based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features of the image. An information processing system equipped with the following features.
2. An information processing system including an electronic device used by a user and an information processing device capable of providing image information to the electronic device, A database that stores images and their characteristics in association with each set of images, An acquisition unit that acquires selected images from the electronic device according to the user's preferences, A text information generation unit that generates text information based on the selected image, A feature generation unit generates features corresponding to predetermined items based on the aforementioned text information and uses them as features of the selected image. A selection unit that, based on a comparison result obtained by comparing the characteristics of the selected image with the characteristics of the image, selects a new image from the images stored in the database that matches the user's preferences. An information processing system equipped with the following features.
3. An information processing system comprising an electronic device used by a user and an information processing device capable of providing the electronic device with recommended images retrieved using selected images chosen according to the user's preferences, A database that stores the characteristics of goods or services provided by a provider and the provider, associating them with each of multiple providers, An acquisition unit that acquires text information indicating requests for the aforementioned recommended images, A feature generation unit generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection unit that selects a provider from among the providers stored in the database that matches the user's preferences based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features relating to the product or the service. An information processing system equipped with the following features.
4. The information processing system according to claim 2, The acquisition unit acquires a plurality of selected images according to the user's preferences from the electronic device, The system further includes a feature information generation unit that generates feature information indicating the user's preferences based on the characteristics of each of the aforementioned multiple selected images. The selection unit selects a new image according to the user's preferences based on the comparison result obtained by comparing the feature information with the features of the image. Information processing system.
5. An information processing system according to claim 1 or 2, The system further includes an information providing unit that transmits display information for displaying the image newly selected by the selection unit to the electronic device, The electronic device includes a control unit that performs display control to display the newly selected image on the display unit based on the display information. Information processing system.
6. An information processing system according to claim 1 or 2, The system further includes an information providing unit that transmits display information for displaying images stored in the database to the electronic device. The acquisition unit acquires from the electronic device an image selected by the user based on user operation, from among the images displayed on the electronic device based on the display information, as the selected image. Information processing system.
7. The information processing system according to claim 6, The aforementioned electronic device is Display control that causes the display unit to display an image stored in the database based on the aforementioned display information, The control unit performs transmission control to transmit edited image information relating to an edited image, which has been edited based on user operation, to the information processing device. The acquisition unit acquires the image corresponding to the edited image information transmitted from the electronic device as the selected image. Information processing system.
8. The information processing system according to claim 6, When the information provision unit receives a re-suggestion request from the electronic device for re-suggestion of a new image other than the newly selected image, it uses the features generated by the feature generation unit to send display information to the electronic device for displaying an image newly selected from the images stored in the database as an image that matches the user's preference. Information processing system.
9. An information processing system according to any one of claims 1 to 3, The aforementioned electronic device is An image information database that stores images acquired by the image acquisition unit, The system includes a control unit that performs display control to display images stored in the image information database on the display unit, and a control unit that performs transmission control to transmit selected image information relating to an image selected from the images displayed on the display unit based on user operation, The acquisition unit acquires the image corresponding to the selected image information transmitted from the electronic device as the selected image. Information processing system.
10. An information processing system according to claim 1 or 3, The feature generation unit generates at least one of the first feature and the second feature using an AI model that extracts features related to what is included in the target image. Information processing system.
11. An information processing device capable of providing recommended images, searched using selected images according to the user's preferences, to the electronic device used by the user, and capable of handling a database that stores images and their characteristics in association with each of multiple images, An acquisition unit that acquires text information indicating requests for the aforementioned recommended images, A feature generation unit generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection unit selects a new image from the images stored in the database that matches the user's preferences based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features of the image. An information processing device equipped with the following features.
12. An information processing device capable of providing image information to an electronic device used by a user, and capable of handling a database that stores images and the characteristics of those images in association with each of a plurality of images, An acquisition unit that acquires selected images from the electronic device according to the user's preferences, A text information generation unit that generates text information based on the selected image, A feature generation unit generates features corresponding to predetermined items based on the aforementioned text information and uses them as features of the selected image. A selection unit that, based on a comparison result obtained by comparing the characteristics of the selected image with the characteristics of the image, selects a new image from the images stored in the database that matches the user's preferences. An information processing device equipped with the following features.
13. An information processing device capable of providing recommended images, retrieved using selected images chosen according to the user's preferences, to an electronic device used by the user, and capable of handling a database that stores the characteristics of goods or services provided by a provider and the provider associated with each of a plurality of providers, An acquisition unit that acquires text information indicating requests for the aforementioned recommended images, A feature generation unit generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection unit that selects a provider from among the providers stored in the database that matches the user's preferences based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features relating to the product or the service. An information processing device equipped with the following features.
14. An information processing method performed by a computer capable of handling a database that stores images and their characteristics in association with each set of images, which can provide recommended images, searched using selected images according to the user's preferences, to the electronic device used by the user, When the computer obtains text information indicating a request for the recommended image, it generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A computer performs a selection process to select a new image from the images stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features of the image. Information processing methods including
15. An information processing method performed by a computer capable of providing image information to an electronic device used by a user, and handling a database that stores images and the characteristics of those images in association with each of a plurality of images, When the computer obtains selected images from the electronic device according to the user's preferences, it performs a text information generation process that generates text information based on the selected images. A computer performs a feature generation process that generates features corresponding to predetermined items based on the text information and uses them as features of the selected image. A computer performs a selection process in which it selects a new image from the images stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the features of the selected image with the features of the image. Information processing methods including
16. An information processing method performed by a computer capable of providing recommended images, which have been searched using selected images chosen according to the user's preferences, to an electronic device used by the user, and which stores a database that associates and stores characteristics of goods or services provided by a provider with the provider for each of a plurality of providers, When the computer obtains text information indicating a request for the recommended image, it generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A computer performs a selection process to select a provider from among the providers stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features relating to the product or service. Information processing methods including
17. A program that can provide recommended images, searched using selected images according to the user's preferences, to the electronic device used by the user, and which is executed on a computer capable of handling a database that stores images and their characteristics in association with each set of images, A feature generation procedure that, upon obtaining text information indicating requests for the recommended image, generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection procedure for selecting a new image from the images stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features of the image. A program that causes a computer to execute something.
18. A program to be executed on a computer capable of providing image information to an electronic device used by a user, and which is capable of handling a database that stores images and their characteristics in association with each of a plurality of images, A text information generation procedure is performed to obtain selected images according to the user's preferences from the electronic device, and to generate text information based on the selected images. A feature generation procedure that generates features corresponding to predetermined items based on the aforementioned text information and uses them as features of the selected image, A selection procedure for selecting a new image from the images stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the characteristics of the selected image with the characteristics of the image. A program that causes a computer to execute something.
19. A program to be executed on a computer capable of providing recommended images, which have been searched using selected images chosen according to the user's preferences, to an electronic device used by the user, and which is capable of handling a database that stores the characteristics of goods or services provided by a provider and the provider associated with each of a plurality of providers, A feature generation procedure that, upon obtaining text information indicating requests for the recommended image, generates a first feature relating to the user's preferences based on the text information, and generates a second feature which is a feature of the selected image relating to what is included in the selected image, based on the selected image. A selection procedure for selecting a provider from among the providers stored in the database that matches the user's preferences, based on a comparison result obtained by comparing the feature information indicating the user's preferences, which is generated based on the first and second features, with the features relating to the product or service. A program that causes a computer to execute something.