Program, method, information processing apparatus, and system
The system uses a multimodal generative AI model to generate explanatory texts from images and extract characteristic words as tags, addressing the accuracy issues of existing image classification tools and creating a high-quality tag cloud.
Patent Information
- Application Number
- JP2024085956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-10
- Estimated Expiration
- 2044-05-28
AI Technical Summary
Existing image classification tools for generating image tags lack accuracy in tagging images, resulting in low-quality tag clouds.
A system utilizing a multimodal generative AI model to generate explanatory texts from images, followed by natural language analysis to extract characteristic words as tags, which are then presented as a tag cloud.
The system achieves a tag cloud composed of highly accurate tags, improving the quality and relevance of image tagging.
Smart Images

Figure 2025179304000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, a method, an information processing device, and a system. [Background technology]
[0002] A user interface often used on social tagging sites is called a tag cloud, which is a list of tags based on the number of tags registered, frequency of appearance, popularity, importance, etc. However, to realize a tag cloud, for example, it is necessary to attach multiple tags to one piece of content. Attaching multiple tags to content requires a lot of work. To address this problem, for example, the technology disclosed in Patent Document 1 uses an artificial intelligence (AI)-based image classification tool to identify specific image features and / or generate image tags. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2022-505237 Summary of the Invention [Problem to be solved by the invention]
[0004] Specifically, image analysis using the aforementioned image classification tool may identify multiple image features, such as a smile, a waitress, a counter, a vending machine, a coffee cup, a cake, a hand, food, a person, a cafe, etc. Images may then be tagged with each of these identified features. However, tagging using an image classification tool does not result in high accuracy of the tags set.
[0005] The objective of the present disclosure is to realize a tag cloud consisting of highly accurate tags. [Means for solving the problem]
[0006] A program to be executed by a computer having a processor and a memory, the program instructs the processor to execute the following steps: accepting input of a plurality of images; inputting, for each of the input images, the image and an instruction statement instructing the multimodal generative AI model to output an explanatory text about the image, causing the multimodal generative AI model to output the explanatory text; analyzing the output explanatory texts and extracting characteristic words as tags; and presenting the extracted tags as a first list. [Effects of the Invention]
[0007] According to the present disclosure, a tag cloud made up of highly accurate tags can be realized as the first list. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an example of the overall configuration of a system 1. FIG. [Figure 2] 2 is a block diagram showing an example of the configuration of a terminal device 10 shown in FIG. 1. FIG. [Figure 3] 2 is a block diagram showing an example of a functional configuration of a server 20 shown in FIG. 1. FIG. [Figure 4] FIG. 10 is a diagram showing the data structure of a first tag cloud. [Figure 5] 4 is a diagram showing the data structure of a user information table 2021 shown in FIG. 3. FIG. [Figure 6] FIG. 4 is a diagram showing the data structure of an image table 2022 shown in FIG. 3. [Figure 7] FIG. 4 is a diagram showing the data structure of a prompt table 2023 shown in FIG. 3. [Figure 8] FIG. 4 is a diagram showing the data structure of an explanation table 2024 shown in FIG. 3. [Figure 9] 10 is a flowchart showing an example of the operation of the server 20 when presenting a first tag cloud to a user. [Figure 10] FIG. 10 is a schematic diagram showing an example of the display screen of display 141 when first tag cloud 40 is presented to the user. [Figure 11] 11 is a schematic diagram showing an example of the display screen of display 141 when a user searches first tag cloud 40 shown in FIG. 10. FIG. [Figure 12] FIG. 10 is a schematic diagram showing another example of the display screen of display 141 when first tag cloud 401 is presented to the user. [Figure 13] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals. The names and functions of the components are also the same. Therefore, detailed descriptions thereof will not be repeated.
[0010] <Summary> The system according to this embodiment has a function for generating tags related to images. For example, for each of a plurality of images input by a user, the system according to this embodiment inputs the image and an instruction statement instructing the output of an explanatory statement related to the image into a multimodal generative AI model (details will be described later). The content of the image input to the multimodal generative AI model is not particularly limited, and may be an image related to a restaurant, apparel store, real estate property, etc. Furthermore, for example, if the image input to the multimodal generative AI model is an image related to a restaurant, izakayas, restaurants, cafes, etc. are also included in the captured objects of the image.
[0011] The system according to this embodiment performs natural language analysis on explanatory text output by a multimodal generative AI model, extracts multiple characteristic words, and sets each of these as tags. The system then accumulates the extracted tags for the multiple explanatory texts. The system presents the accumulated tags to the user as a list. This list becomes a first tag cloud. This makes it possible to realize a first tag cloud composed of highly accurate tags.
[0012] <1 Overall system configuration> Fig. 1 is a block diagram showing an example of the overall configuration of system 1. System 1 shown in Fig. 1 includes, for example, a terminal device 10, a server 20, and a multimodal generation AI system 30 (hereinafter referred to as "MAI system 30"). Terminal device 10, server 20, and MAI system 30 are communicatively connected via, for example, a network 80.
[0013] 1 shows an example in which system 1 includes two terminal devices 10, but the number of terminal devices 10 included in system 1 may be less than three or may be three or more. While FIG. 1 shows an example in which system 1 includes one MAI system 30, the number of MAI systems 30 included in system 1 may be two or more.
[0014] 1 shows an example in which MAI system 30 is independent from server 20, server 20 may include the functionality of MAI system 30. In other words, server 20 may store a multimodal generative AI model.
[0015] In this embodiment, a collection of multiple devices may be considered as one server. The way in which the multiple functions required to realize the server 20 according to this embodiment are allocated to one or more pieces of hardware can be determined appropriately depending on the processing capacity of each piece of hardware and / or the specifications required for the server 20.
[0016] 1 is an information processing device operated by a user who uses a first tag cloud (first list). The terminal device 10 is realized, for example, by a desktop personal computer (PC), a laptop PC, etc. The terminal device 10 may also be realized by a mobile terminal such as a smartphone or a tablet.
[0017] The terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage 16, and a processor 19. The input device 13 is a device for receiving input operations from a user (for example, a touch panel, a touch pad, a pointing device such as a mouse, a keyboard, etc.). The output device 14 is a device for presenting information to a user (a display, a speaker, etc.).
[0018] The server 20 is, for example, an information processing device that provides a first tag cloud service using a multimodal generative AI model. The server 20 transmits a plurality of images input by a user via the terminal device 10 to the MAI system 30. The server 20 transmits a command input by a user via the terminal device 10 to the MAI system 30.
[0019] The instruction sentence is a sentence that instructs the output of an explanatory sentence (hereinafter referred to as "explanation sentence") related to the image input by the user. For example, the instruction sentence may have different instruction contents input for each of the multiple images, or the same instruction contents may be input for some or all of the multiple images. Note that the instruction sentence does not have to be input to the server 20 via the terminal device 10. For example, one or more instruction sentences defaulted according to the type of imaged object shown in each of the multiple images may be stored in the prompt table 2023 in advance. For example, when an image is input, the server 20 selects an instruction sentence stored in the prompt table 2023 in accordance with a predetermined rule.
[0020] The server 20 outputs the description to the MAI system 30. For example, if the captured object shown in the image is a restaurant, the description would be a promotional message for the restaurant. The server 20 analyzes the description output from the MAI system 30 and extracts multiple tags. The server 20 transmits the extracted multiple tags as a first tag cloud to the terminal device 10, which presents the first tag cloud to the user via the output device 14.
[0021] There are no particular limitations on the manner in which the first tag cloud is presented to the user. Server 20 may design the text color, text size, font, display direction, and layout of the tags that make up the first tag cloud in various ways, taking into account user convenience, etc.
[0022] The server 20 is, for example, an information processing device realized by a computer connected to a network 80. As shown in Fig. 1, the server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29. The input / output IF 23 functions as an input device for receiving input operations from a user and as an interface for an output device for outputting information to the user.
[0023] The MAI system 30 includes a multimodal generative AI model. The multimodal generative AI model is a large-scale language model constructed by deep learning that collects two or more types of information, such as text, audio, images, and video, and processes them in an integrated manner. Examples of multimodal generative AI models include "OpenAI ChatGPT GTP-4 (registered trademark)" and "Google Gemini."
[0024] The MAI system 30 inputs the image and instruction received from the server 20 into the multimodal generative AI model and causes the multimodal generative AI model to output an explanatory text. The MAI system 30 transmits the explanatory text output by the multimodal generative AI model to the server 20.
[0025] Each information processing device is configured by a computer 90 (see FIG. 11) equipped with an arithmetic unit and a storage device. The basic hardware configuration of the computer 90 and the basic functional configuration of the computer 90 realized by the basic hardware configuration will be described later. For the terminal device 10, the server 20, and the MAI system 30, descriptions that overlap with the basic hardware configuration and basic functional configuration of the computer 90 will be omitted.
[0026] <1.1 Terminal device configuration> Fig. 2 is a block diagram showing an example of the configuration of the terminal device 10 shown in Fig. 1. As shown in Fig. 2, the terminal device 10 includes a communication unit 120, an input device 13, an output device 14, a storage unit 180, and a control unit 190. The blocks included in the terminal device 10 are electrically connected by, for example, a bus. The terminal device 10 may also include an audio processing unit, a microphone, a speaker, a camera, a position information sensor, or a combination of at least two of these.
[0027] The communication unit 120 performs processing such as modulation and demodulation for the terminal device 10 to communicate with other devices. The communication unit 120 performs transmission processing on signals generated by the control unit 190 and transmits the signals to the outside (for example, the server 20). The communication unit 120 performs reception processing on signals received from the outside and outputs the signals to the control unit 190.
[0028] The input device 13 is a device for a user operating the terminal device 10 to input instructions or information. The input device 13 is realized, for example, by a touch-sensitive device 131 that inputs instructions by touching the operation surface. If the terminal device 10 is a PC or the like, the input device 13 may be realized by a reader, keyboard, mouse, or the like. The input device 13 converts instructions input by the user into electrical signals and outputs them to the control unit 190. The input device 13 may include, for example, a receiving port that receives electrical signals input from an external input device.
[0029] The output device 14 is a device for presenting information to a user operating the terminal device 10. The output device 14 is realized, for example, by a display 141. The display 141 displays various information according to the control of the control unit 190. The display 141 is realized, for example, by an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0030] The storage unit 180 is realized by, for example, the memory 15 and the storage 16, and stores data and programs used by the terminal device 10. The storage unit 180 stores, for example, user information 181. The user information 181 includes, for example, various information related to the user who uses the terminal device 10. The various information related to the user includes, for example, the user's name, age, address, date of birth, and contact information.
[0031] The control unit 190 is realized by the processor 19 reading a program stored in the storage unit 180 and executing instructions included in the program. The control unit 190 controls the operation of the terminal device 10. The control unit 190 performs the functions of an operation reception unit 191, a transmission / reception unit 192, and a presentation control unit 193 by operating in accordance with the program.
[0032] The operation reception unit 191 performs processing for receiving instructions or information input from the input device 13. Specifically, the operation reception unit 191 receives instructions or information input from the touch-sensitive device 131.
[0033] The transmitting / receiving unit 192 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol. Specifically, the transmitting / receiving unit 192 transmits instructions or information input by the user to the server 20. The transmitting / receiving unit 192 receives information transmitted from the server 20.
[0034] The presentation control unit 193 controls the output device 14 to present to the user the information transmitted from the server 20. Specifically, the presentation control unit 194 causes the display 141 to display the first tag cloud transmitted from the server 20.
[0035] <1.2 Functional configuration of the server> Fig. 3 is a block diagram showing an example of the functional configuration of the server 20 shown in Fig. 1. As shown in Fig. 3, the server 20 performs the functions of a communication unit 201, a storage unit 202, and a control unit 203.
[0036] The communication unit 201 performs processing for communication between the server 20 and external devices. The storage unit 202 includes, for example, a user information table 2021, an image table 2022, a prompt table 2023, and an explanation table 2024.
[0037] The tables stored in the storage unit 202 are not limited to these. For example, the image table 2022 does not have to be stored in the storage unit 202. The service using the first tag cloud may differ depending on the content of the image. Therefore, the images may be stored for each server managing the service. For example, the image table 2022 may be stored in a server related to a service for sharing information among multiple users. For example, the storage unit 202 may have a first tag cloud table 2025 as shown in FIGS. 3 and 4. The first tag cloud table 2025 is a table that stores the first tag cloud, and as shown in FIG. 4, has columns such as tag data and target, with a tag cloud ID as a key. The tag cloud ID is an item that stores an identifier for uniquely identifying the first tag cloud. The tag data is an item that stores data related to multiple tags that make up the first tag cloud. The target is an item that stores the captured object shown in the image.
[0038] The user information table 2021 is a table that stores information about users. The image table 2022 is a table that stores information about a plurality of images input by the user. The prompt table 2023 is a table that stores information about instruction sentences that are input sentences to the MAI system 30. The explanation sentence table 2024 is a table that stores information about explanation sentences that are output sentences from the MAI system 30. Details of these tables will be described later.
[0039] The control unit 203 is realized by the processor 29 reading a program stored in the storage unit 202 and executing instructions included in the program. The program includes an application such as a web browser application. The program includes a programming language such as JavaScript (registered trademark) that is executed on the web browser application stored in the terminal device 10. The control unit 203 operates in accordance with the program to fulfill the functions of a reception control module 2031, a transmission control module 2032, a tag generation module 2033, and a presentation control module 2035.
[0040] The reception control module 2031 controls the process of the server 20 receiving signals from external devices in accordance with a communication protocol. The transmission control module 2032 controls the process of the server 20 transmitting signals to external devices in accordance with a communication protocol.
[0041] The tag generation module 2033 executes a process of generating multiple tags using the explanatory text output from the MAI system 30. Specifically, the tag generation module 2033 analyzes the explanatory text output from the MAI system 30. For example, the tag generation module 2033 performs natural language analysis on the explanatory text. The natural language analysis may utilize existing technology or may be performed using a predetermined trained model. Based on the analysis results, the tag generation module 2033 extracts multiple characteristic words from the explanatory text and sets each of the extracted multiple words as a tag.
[0042] Here, a "characteristic word" refers to a word that expresses a meaning that allows a service user of the first tag cloud to easily recall the imaged object shown in the image. The service user who serves as the criterion for determining whether a word is a "characteristic word" is assumed to be someone with an average level of knowledge about the imaged object. The tag generation module 2033 determines whether a word included in the description is a characteristic word based on, for example, dictionary information pre-stored in the storage unit 202. The dictionary information may be stored for each predetermined field. The dictionary information may be updated as necessary. Furthermore, the dictionary information may learn new information in response to user operations.
[0043] The presentation control module 2035 controls the process of presenting information to a user. For example, the presentation control module 2035 controls the process of presenting a plurality of tags to a user by configuring the plurality of tags generated by the tag generation module 2033 as a first tag cloud.
[0044] <2 Data Structure> 5 to 8 are diagrams showing the data structure of each table stored in the server 20. Note that Figures 5 to 8 are merely examples and do not exclude data that is not listed. Furthermore, even data listed in the same table may be stored in separate storage areas in the storage unit 202.
[0045] Fig. 5 is a diagram showing the data structure of the user information table 2021. The user information table 2021 shown in Fig. 5 is a table having columns of name, age, sex, date of birth, and contact information, with the user ID as a key.
[0046] The user ID is an item that stores an identifier for uniquely identifying a user. The name is an item that stores the user's name. The age is an item that stores the user's age. The gender is an item that stores the user's gender. The date of birth is an item that stores the user's date of birth. The contact information is an item that stores the contact information (e.g., telephone number, email address, etc.) of the terminal device 10 that the user has.
[0047] Fig. 6 is a diagram showing the data structure of image table 2022. Image table 2022 shown in Fig. 6 is a table having columns for user ID, date and time, subject, and image, with image ID as a key. The columns included in image table 2022 are not limited to the example in Fig. 6. Image table 2022 may also include columns for storing information about facilities associated with images, such as URLs, business hours, addresses, contact information, and recommended information.
[0048] The image ID is an item that stores an identifier for uniquely identifying an image. The user ID is an item that stores an identifier for uniquely identifying a person who inputs an image. The user ID stored in the image table 2022 corresponds to, for example, the user ID stored in the user information table 2021. Note that, instead of the user ID, the name of the person who inputs the image may be stored. The date and time is an item that stores the date and time when the image was captured. Note that, the date and time may also be an item that stores the date and time when the image was input.
[0049] The "object" is an item that stores the captured object shown in the image. The captured object includes, for example, a restaurant, an apparel store, and a real estate property. Also, for example, if the captured object is a restaurant, izakayas, restaurants, cafes, etc. are also included in the captured object. The "image" is an item that stores image data input by the user. If the image data is related to a restaurant or an apparel store, the "image" item stores, for example, an image showing the interior of the restaurant or the apparel store. If the image data is related to a real estate property, the "image" item stores, for example, an image showing the exterior of the real estate property.
[0050] Fig. 7 is a diagram showing the data structure of the prompt table 2023. The prompt table 2023 shown in Fig. 7 is a table having prompt data and target columns with a prompt ID as a key.
[0051] The prompt ID is an item that stores an identifier for uniquely identifying an instruction sentence. The prompt data is an item that stores data related to an instruction sentence. The data related to an instruction sentence is, for example, text information indicating the content of the instruction sentence. The text information indicating the content of the instruction sentence is, for example, an input sentence for causing the MAI system 30 to create an explanatory sentence, and includes, for example, the following content: "The image you entered was taken at ... (restaurant). Please create a description for this image that includes the following points: ·(Restaurant) History Current customer demographics and store visits · (Restaurant) interior and exterior atmosphere ·menu Location
[0052] The item "prompt data" may store a reference (path) to a prompt data file located elsewhere.
[0053] The object is an item that stores the imaged object shown in the image. The objects stored in the prompt table 2023 correspond to the objects stored in the image table 2022, for example. In other words, prompt data is stored in advance for each imaged object, for example.
[0054] Fig. 8 is a diagram showing the data structure of the description table 2024. The description table 2024 shown in Fig. 8 is a table having columns of description data and image ID, with a description ID as a key.
[0055] The description ID is an item that stores an identifier for uniquely identifying a description. The description data is an item that stores data related to the description. The data related to the description is, for example, text information indicating the content of the description. The text information indicating the content of the description is a sentence that serves as the basis for extracting multiple tags that make up the first tag cloud, and includes, for example, the following content: "...(the restaurant) is a new cafe that opened a month ago. Located on the riverside that runs through the center of the downtown area, you can enjoy a relaxing time in a calm atmosphere while admiring the scenery around the river. Its easy access makes it popular with a wide range of customers, from young people to seniors. In addition to cafe menu items such as desserts, it also has a full lunch menu, making it suitable for a variety of occasions."
[0056] The item "explanatory data" may store reference information (path) to a prompt data file located elsewhere.
[0057] The image ID is an item that stores an identifier for uniquely identifying an image, similar to the image IDs stored in the image table 2022 and the prompt table 2023, respectively.
[0058] <3 operations> With reference to FIG. 9, the operation of server 20 when presenting the first tag cloud to a user will be described. FIG. 9 is a flowchart showing an example of the operation of server 20 when presenting the first tag cloud to a user. In the description of FIG. 9, it is assumed that the types of image capture objects shown in each of the multiple images input by the user are all restaurants. The image capture objects may be specified by the user when inputting the images. Alternatively, the image capture objects may be recognized by performing image analysis on the input images. Alternatively, the image capture objects may be recognized based on an information site that registers information related to images. For example, when registering information related to images on a gourmet site, the image capture objects are recognized as "restaurants." Note that the number of images input by the user does not have to be multiple, but may be one.
[0059] In step S11, the server 20 accepts input of a plurality of images. Specifically, the user operates the input device 13 to input a plurality of images to the terminal device 10. The operation acceptance unit 191 accepts the input of a plurality of images by the user. The transmission / reception unit 192 transmits the plurality of images accepted by the operation acceptance unit 191 to the server 20. The reception control module 2031 receives the plurality of images sent from the terminal device 10. In other words, the server 20 accepts the input of a plurality of images by the user.
[0060] In step S12, the server 20 accepts the input of an instruction statement. Specifically, the user operates the input device 13 to input the instruction statement into the terminal device 10. The operation acceptance unit 191 accepts the input of the instruction statement by the user. The transmission / reception unit 192 transmits the instruction statement accepted by the operation acceptance unit 191 to the server 20. The reception control module 2031 receives the instruction statement transmitted from the terminal device 10. In other words, the server 20 accepts the input of the instruction statement by the user. Note that if the instruction statement is not input by the user but is pre-stored in the prompt table 2023, the system 1 does not need to execute the processing of step S12.
[0061] In step S13, the server 20 causes the multimodal generative AI model to output explanatory text corresponding to each of the multiple images. Specifically, the transmission control module 2032 transmits the multiple images received by the reception control module 2031 to the MAI system 30. Basically, the transmission control module 2032 sequentially transmits each image to the MAI system 30 as the reception control module 2031 receives it. However, the transmission control module 2032 may transmit all of the multiple images to the MAI system 30 at once, for example. The transmission control module 2032 transmits the instruction text received by the reception control module 2031 to the MAI system 30.
[0062] When receiving input of instruction sentences corresponding to a plurality of images, each time the reception control module 2031 receives an instruction sentence, the transmission control module 2032 transmits the instruction sentence together with the corresponding image to the MAI system 30. However, the timing of transmitting the instruction sentence and the timing of transmitting the image may be different.
[0063] When input of a command from the user is not required, the transmission control module 2032 reads prompt data from the prompt table 2023 based on the item "target" corresponding to the imaged object shown in the input image. The transmission control module 2032 transmits the read prompt data to the MAI system 30 as a command.
[0064] The MAI system 30 inputs the multiple images and instruction sentences received from the server 20 into the multimodal generative AI model. The MAI system 30 causes the multimodal generative AI model to output explanatory sentences corresponding to each of the multiple images based on the multiple images and instruction sentences input into the multimodal generative AI model. The MAI system 30 transmits the multiple explanatory sentences output from the multimodal generative AI model to the server 20. The reception control module 2031 receives the multiple explanatory sentences transmitted from the MAI system 30. In other words, the server 20 causes the multimodal generative AI model to output the multiple explanatory sentences.
[0065] In step S14, the server 20 extracts a plurality of tags. Specifically, the tag generation module 2033 analyzes the received plurality of descriptions and extracts a plurality of characteristic words as tags. The tag generation module 2033 associates the extracted plurality of tags with images and stores them in the first tag cloud table 2025. The server 20 generates a tag for each image to be registered and accumulates the tags in the first tag cloud table 2025.
[0066] In step S15, the server 20 presents the first tag cloud to the user. Specifically, for example, when the user requests display of the first tag cloud, that is, when the user requests a search for restaurants, the presentation control module 2035 reads tag data from the first tag cloud table 2025 by referring to the item "target." The presentation control module 2035 configures the read tag data as a first tag cloud and displays the configured first tag cloud on the terminal device 10. That is, the transmission control module 2032 transmits information for displaying the first tag cloud configured by the presentation control module 2035 to the terminal device 10. That is, the server 20 presents the first tag cloud to the user.
[0067] Note that the presentation control unit 193 may configure a plurality of tags as the first tag cloud instead of the presentation control module 2035. In this case, the transmission control module 2032 transmits the tag data read from the first tag cloud table 2025 to the terminal device 10.
[0068] Transmitting / receiving unit 192 receives information for displaying the first tag cloud transmitted from server 20. Presentation control unit 193 causes display 141 to display the first tag cloud based on the received information for displaying the first tag cloud. Note that step S15 does not necessarily have to be performed after step S14. Server 20 may execute step S15 in response to a request from the user.
[0069] <4 Screen example> Fig. 10 is a schematic diagram showing an example of the display screen of display 141 when first tag cloud 40 is presented to the user. That is, Fig. 10 is an example of a search screen using the first tag cloud. Fig. 10 illustrates an example in which the type of imaged object shown in each of the multiple images input by the user is a cafe (restaurant). The same applies to Figs. 11 and 12 described below.
[0070] The display screen shown in FIG. 10 displays a plurality of tags 41 reminiscent of a cafe in a list format, that is, in a list. The plurality of tags 41 are based on, for example, tags registered for the same image capture object. This block of the plurality of tags 41 displayed in a list becomes the first tag cloud 40. Specifically, the first tag cloud 40 shown in FIG. 10 is made up of three rows of records, and each row contains a different number of types of tags 41. Furthermore, in the first tag cloud 40, the text color, text size, font, and display direction of all tags 41 are the same. For example, the display direction of the plurality of tags 41 is all horizontal (left to right as you face the page).
[0071] 10, first tag cloud 40 is displayed in area 1411 on the display screen, but first tag cloud 40 may be displayed in any area on the display screen. This also applies to the examples of FIGS. 11 and 12 described below. The placement of first tag cloud 40 on the display screen is controlled by, for example, presentation control module 2035 or presentation control unit 193.
[0072] The user operates the first tag cloud 40 displayed on the display screen shown in Fig. 10, for example. That is, the user selects one of the tags 41 included in the first tag cloud 40. The user may select one tag 41 included in the first tag cloud 40, or may select multiple tags 41. "The user's operation of selecting tags 41" specifically refers to, for example, an operation in which the user selects one or more desired tags 41 from the multiple tags 41 that make up the first tag cloud 40.
[0073] Here, the control unit 203 functions as an image-related information generation module. In this embodiment, the image-related information is, for example, information about an image that includes one or more selected tags in its description. Specifically, the image-related information includes, for example, an image that includes one or more selected tags in its description, a description of the image, information about the store represented by the image, or at least any combination of these. The information about the store represented by the image includes the store's URL, business hours, address, contact information, recommended information, etc.
[0074] When the image-related information generation module receives a tag selection from the user, it generates image-related information based on the selected tag. The image-related information may be stored, for example, in an image-related information table (not shown) stored in the storage unit 202. The presentation control module 2035 presents the image-related information to the user. For example, the transmission control module 2032 transmits information for displaying the image-related information to the terminal device 10. When the information for displaying the image-related information is received, the presentation control unit 193 transitions the display screen shown in FIG. 10 to, for example, the display screen shown in FIG. 11.
[0075] Fig. 11 is a schematic diagram showing an example of a display screen of display 141 when a user operates first tag cloud 40 shown in Fig. 10. On the display screen shown in Fig. 11, tags 41 selected in first tag cloud 40 are displayed in area 1412. Also, on the display screen shown in Fig. 11, image-related information 50 generated based on tags 41 selected in the first tag cloud is displayed in area 1413. In the example of Fig. 11, as a search result using the first tag cloud, articles introducing several cafes that are expected to meet the user's desired conditions are displayed in a list format as image-related information 50.
[0076] However, the display mode of the search results of the first tag cloud is not limited to the example of FIG. 11 , and various display modes can be adopted. For example, in the example of FIG. 11 , before the image-related information 50 is displayed, only the number of pieces of image-related information 50 generated by the selected tag may be displayed. The user checks the number of pieces displayed and determines whether to display the image-related information 50 or select additional tags 41. For example, if the number of pieces displayed is greater than expected, the user selects additional tags 41 from the first tag cloud 40. Furthermore, if the number of pieces of image-related information 50 is tolerable, the user inputs an instruction to display the image-related information 50 from the input device 13. Note that, for the sake of simplicity of explanation and illustration, two pieces of image-related information 50 are displayed in the example of FIG. 11 , but the number of pieces of image-related information to be displayed is not limited to two.
[0077] <5 Summary> In this embodiment, the operation acceptance unit 191 and the reception control module 2031 accept input of multiple images. The MAI system 30 inputs the image and instruction text for each of the multiple images to the multimodal generative AI model. The MAI system 30 causes the multimodal generative AI model to output explanatory text corresponding to each of the multiple images based on the multiple images and instruction text input to the multimodal generative AI model. The tag generation module 2033 analyzes the explanatory text and extracts multiple characteristic words as tags. The presentation control module 2035 organizes the multiple tags into a first tag cloud and presents it to the user.
[0078] As described above, in this embodiment, tags are generated using explanatory text obtained by inputting multiple images and instructional text into a multimodal generative AI model. Therefore, the accuracy of tags is improved compared to when tags are generated using an AI-based image classification tool. Therefore, this embodiment can realize a first tag cloud composed of highly accurate tags.
[0079] <6. First Modification> The system 1 may, for example, analyze the extracted tags and present the first tag cloud in a manner according to the analysis results (first modification). The analysis of the tags may, for example, analyze the number of extracted tags, the frequency of appearance, the importance of the tags, or a combination of at least two of these.
[0080] The analysis of the multiple tags is realized, for example, by the control unit 190 or the control unit 203 functioning as a tag analysis module (not shown). For example, the tag analysis module analyzes the number of appearances and frequency of appearances of tags generated by the tag generation module 2033. The tag analysis module also analyzes the importance of tags using information such as the frequency of appearance of the tags. The presentation of the first tag cloud in a manner corresponding to the analysis results is realized, for example, by the control unit 190 functioning as the presentation control unit 193. Alternatively, the presentation of the first tag cloud in a manner corresponding to the analysis results is realized by the control unit 203 functioning as the presentation control module 2035. The analysis results of the multiple tags may be stored, for example, in an analysis result table (not shown) included in the storage unit 202, or may be stored in the first tag cloud table 2025.
[0081] According to the first modification, the analysis results for each of the multiple tags are reflected in the presentation of the first tag cloud. This allows the user to select tags from the first tag cloud that are more likely to lead to the content, service, etc. that the user desires. In other words, a first tag cloud composed of tags with higher accuracy can be realized.
[0082] In the first modification, the presentation mode of the first tag cloud according to the analysis result may be a mode according to the number of extracted tags, the frequency of appearance, the importance, or a combination of at least two of these. This configuration improves the accuracy of tag selection by the user compared to presentation modes according to other analysis results.
[0083] In the first modified example, the presentation mode of the first tag cloud may be the size, color, font, or display orientation of the characters of the words constituting the tags, or a combination of at least two of these. This configuration improves the accuracy of tag selection by the user compared to other presentation modes. An example of this configuration will be described below with reference to the example in FIG. 12. FIG. 12 is a schematic diagram showing another example of the display screen of display 141 when first tag cloud 401 is presented to the user.
[0084] 12, a plurality of tags 411 reminiscent of cafes are displayed in list format, that is, in a list. A block of the plurality of tags 411 displayed in this list becomes first tag cloud 401. Specifically, first tag cloud 401 is not composed of columns and records, and individual tags 411 are arranged randomly within the block of first tag cloud 401. Furthermore, tags 411 whose display direction is vertical (up and down as you face the page) and tags 411 whose display direction is horizontal are mixed within the block of first tag cloud 401.
[0085] In the example of FIG. 12 , the size and color of the characters of the words constituting the tags 411 are divided into three levels according to the number of extracted tags 411, the frequency of appearance, the importance, or a combination of at least two of these. Regarding the size of the characters, the characters of the words constituting the tags 411 with high-level features are the largest. Furthermore, the characters of the words constituting the tags 411 with low-level features are the smallest. Furthermore, the size of the characters of the words constituting the tags 411 with medium-level features is intermediate between the tags 411 with high-level features and the tags 411 with low-level features. Regarding the color of the characters, although not shown, the characters of the words constituting the tags 411 with high-level features are red. Furthermore, the characters of the words constituting the tags 411 with low-level features are blue. Furthermore, the characters of the words constituting the tags 411 with medium-level features are yellow.
[0086] The size and color classification of the characters of the words constituting the tag 411 do not have to be three levels, but may be two levels or four or more levels. In the example of FIG. 12, the font of the characters of the words constituting the tag 411 is all the same, and the thickness of the characters increases or decreases in proportion to the size of the characters, but this is not limiting. For example, the type of font may be different depending on the level of the feature possessed by the tag 411. For example, the size of the characters may be the same regardless of the level of the feature possessed by the tag 411, while the thickness of the characters may be different depending on the level.
[0087] Furthermore, the display directions of tags in first tag cloud 401 do not have to be limited to two directions, vertical and horizontal. For example, instead of these directions, the tag display directions may be a first direction tilted at a predetermined angle from vertical to horizontal, and a second direction tilted at a predetermined angle from horizontal to vertical. Alternatively, in addition to the two directions, vertical and horizontal, the tag display directions may be multiple directions including at least one of the first direction and the second direction. Two directions may be used, and first tag cloud 401 is displayed in area 1411 of the display screen, as in the example of FIG. 10.
[0088] <7 Second Modification> For example, the system 1 may accept a tag selection from the first tag cloud, repeat the process of presenting a second tag cloud (second list), and present image-related information (second modified example). The second tag cloud is a tag cloud in which the cluster to which the selected tag belongs is reconfigured. A cluster is each group obtained by classifying and grouping multiple tags that make up the first tag cloud into tags with similar semantics. The image-related information is information about images that include the selected tag in their description. In this case, the "selected tag" refers to the tag selected for cluster extraction.
[0089] Specifically, the image-related information may include, for example, an image containing the selected tag in its description, a description of the image, information about the store represented by the image, or at least any combination thereof. The information about the store represented by the image may include the store's URL, business hours, address, contact information, recommendations, etc.
[0090] The user operates the first tag cloud, i.e., the user selects one of the tags included in the first tag cloud, in the same way as when operating first tag cloud 40 displayed on the display screen shown in Fig. 10. Here, control unit 203 functions as a cluster extraction module. Upon receiving a tag selection from the user, the cluster extraction module extracts a cluster to which the selected tag belongs. The cluster may be stored, for example, in a cluster table (not shown) stored in storage unit 202.
[0091] The presentation control module 2035 presents the second tag cloud to the user. Specifically, the presentation control module 2035 reads, for example, tag data including tags belonging to the extracted cluster from the first tag cloud table 2025. The presentation control module 2035 configures the read tag data as a second tag cloud and presents the configured second tag cloud to the user. That is, the transmission control module 2032 transmits information for displaying the second tag cloud configured by the presentation control module 2035 to the terminal device 10. The transmission / reception unit 192 receives information for displaying the second tag cloud transmitted from the server 20. The presentation control unit 193 displays the second tag cloud on the display 141 based on the received information for displaying the second tag cloud.
[0092] The presentation control module 2035 repeats the process of presenting the second tag cloud to the user. Specifically, for example, after a specific second tag cloud is displayed on the display 141, the operation receiving unit 191 receives a selection of a tag other than the tags constituting the specific second tag cloud. When the cluster extraction module receives the selection of the other tag from the user, it extracts a new cluster to which the other tag belongs from the specific second tag cloud and stores the new cluster in the cluster table. When the user requests the display of a new second tag cloud, the presentation control module 2035 reads other tag data from the cluster table, constructs a new second tag cloud, and presents the new second tag cloud to the user. Repeating this series of processes corresponds to "repeating the process of presenting a second tag cloud." The second tag cloud may be stored, for example, in a second tag cloud table (not shown) in the storage unit 202.
[0093] The presentation control module 2035 presents the image-related information to the user. Specifically, the transmission control module 2032 transmits, for example, information for displaying the image-related information generated by the presentation control module 2035 to the terminal device 10. The presentation control unit 193 displays a second tag cloud on the display 141 based on the received information for displaying the image-related information. Note that, for example, only information related to an image whose description includes the initially selected tag may be displayed on the display 141 as the image-related information. For example, each time a different tag is selected as the image-related information, new image-related information related to an image whose description includes the different tag may be displayed on the display 141. For example, the second tag cloud and the image-related information may be displayed simultaneously on the display 141, or may be displayed at different times. Furthermore, the image-related information may be stored, for example, in an image-related information table (not shown) stored in the storage unit 202.
[0094] According to the second modification, each time the second tag cloud is repeatedly presented to the user, the user can receive a tag cloud service in which the similarity between tags increases. This improves the accuracy of tag selection by the user. In addition, the user can refer to image-related information when selecting a tag from the second tag cloud. This improves the user's convenience when using the second tag cloud.
[0095] In a second modified example, the system 1 may present the number of pieces of image-related information that include the selected tag in their descriptions. Counting the number of pieces of image-related information is achieved, for example, by the control unit 190 or the control unit 203 functioning as a number counting module (not shown). Presentation of the number of pieces of image-related information is achieved, for example, by the control unit 190 functioning as the presentation control unit 193. Alternatively, presentation of the number of pieces of image-related information can also be achieved by the control unit 203 functioning as the presentation control module 2035.
[0096] With this configuration, when a user selects a tag from the second tag cloud, the user can refer to not only the image-related information but also the number of pieces of image-related information related to the second tag cloud. This allows the system 1 to narrow down the number of pieces of image-related information presented to the user each time the process of accepting tag selection from the second tag cloud is repeated. This makes it easier for the user to find the image-related information they desire. This further improves the user's convenience when using the second tag cloud.
[0097] In the second modified example, the system 1 may receive a user's request for the presentation of image-related information and present the image-related information to the user based on the received request. The request may be received, for example, by the control unit 190 functioning as the operation receiving unit 191. Alternatively, the control unit 203 may function as the reception control module 2031. Specifically, for example, assume that the number of image-related information items is displayed on the display 141 after a tag is selected for the second tag cloud. When the user determines that the displayed number of image-related information items is available for review, the user requests the operation receiving unit 191 to present the image-related information. Upon receiving a tag selection from the user, the image-related information generation module generates image-related information based on the selected tag. The presentation control module 2035 presents the user with image-related information whose description includes multiple tags selected by the user. According to this configuration, when the user selects a tag from the second tag cloud, the image-related information can be presented at the user's desired timing. This further improves the user's convenience when using the second tag cloud.
[0098] <8. Third Modification> For example, the system 1 may input, for each of a plurality of input images, text information, the image, and an instruction statement instructing the output of an explanatory statement that takes into account the content of the text information into account into the multimodal generative AI model (third variant). The text information is character information related to the image. Specifically, the "explanatory statement that takes into account the content of the text information" refers to a statement that explains the content of the text information and the image in a comprehensive manner. The acquisition of the text information is realized, for example, by the control unit 190 fulfilling the function of the operation receiving unit 191. Alternatively, it can also be realized by the control unit 203 fulfilling the function of the reception control module 2031.
[0099] According to the third modification, multiple tags are extracted based on explanatory text that takes into account the content of the text information, which further improves the accuracy of the tags. As a result, a first tag cloud made up of tags with higher accuracy can be realized.
[0100] In the third modification, the text information may include, for example, information about a facility related to the image. Also, for example, the text information may include word-of-mouth reviews by users of the facility related to the image. If the captured object shown in the image is a restaurant, an example of a "facility related to the image" is a commercial facility such as a department store that includes the restaurant as a tenant.
[0101] <9 Basic computer hardware configuration> 13 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 91, a main memory device 92, an auxiliary memory device 93, and a communication IF (interface) 99. These components are electrically connected to one another via a bus.
[0102] The processor 91 is hardware for executing an instruction set written in a program. The processor 91 is composed of an arithmetic unit, registers, peripheral circuits, etc. The main memory device 92 is for temporarily storing programs and data processed by the programs, etc., and is, for example, a volatile memory such as a DRAM (Dynamic Random Access Memory). The auxiliary memory device 93 is a memory device for saving data and programs, and is, for example, a flash memory, an HDD (Hard Disc Drive), an optical magnetic disk, a CD-ROM, a DVD-ROM, a semiconductor memory, etc. The communication IF 99 is an interface for inputting and outputting signals for communicating with other computers via a network using wired or wireless communication standards.
[0103] A network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, networks include 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, networks also include those that are directly connected using a USB (Universal Serial Bus) cable, etc.
[0104] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 90 and interconnecting them via a network. In this way, the computer 90 is a concept that includes not only a computer in which each piece of hardware is housed in a single housing or the like, but also a virtualized computer system.
[0105] <Basic functional configuration of computer 90> The functional configuration of a computer realized by the basic hardware configuration of a computer 90 will be described. The computer performs the functions of a control unit, a storage unit, and a communication unit. Note that each function of the computer 90 can be performed by distributing all or part of the functions among multiple computers 90 connected to each other via a network. In this way, the concept of a computer 90 includes not only a single computer but also a virtualized computer system.
[0106] The memory unit is realized by the main memory device 92 and the auxiliary memory device 93. The memory unit stores data, various programs, and various databases. The control unit is realized by the processor 91 reading out various programs stored in the auxiliary memory device 93, expanding them in the main memory device 92, and executing processing in accordance with the programs. The control unit can perform various information processing functions depending on the type of program. In this way, the computer is realized as an information processing device that processes information.
[0107] The control unit can also cause the processor 91 to allocate a storage area corresponding to the storage unit in the main storage unit 92 or the auxiliary storage unit 93 in accordance with the programs. Furthermore, the control unit can cause the processor 91 to add, update, and delete data stored in the storage unit in accordance with various programs.
[0108] A database refers to a relational database, which manages data sets called tables, which are structured by rows and columns, by associating them with each other. In a database, a table is called a table, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables can be set and associated.
[0109] Typically, each table constituting a database has a column set as a key for uniquely identifying a record, but setting a key to a column is not essential. The control unit can cause the processor 91 to add, update, or delete records in a specific table stored in the storage unit in accordance with various programs.
[0110] The communication unit is realized by the communication IF 99. The communication unit has the function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and output the information to the control unit. The control unit can cause the processor 91 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.
[0111] Although several embodiments of the present disclosure have been described above, these embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included in the scope of the inventions and their equivalents as defined in the claims, as well as in the scope and spirit of the inventions.
[0112] <Additional Notes> The matters described in the above embodiments will be supplemented below. (Appendix 1) A program to be executed by a computer having a processor and a memory, the program causing the processor to execute the following steps: accepting input of a plurality of images; for each of the input images, inputting the image and an instruction statement instructing the output of an explanatory text about the image into a multimodal generative AI model and outputting the explanatory text from the multimodal generative AI model; analyzing the output explanatory texts and extracting a plurality of characteristic words as tags; and presenting the extracted tags as a first list. (Appendix 2) A program described in Appendix 1, which causes the processor to execute a step of analyzing the extracted multiple tags, and in a step of presenting the first list, presents the first list in a manner corresponding to the analysis results of the analyzing step. (Appendix 3) A program described in Appendix 2, wherein in the step of presenting the first list, the manner in accordance with the analysis results is a manner in accordance with the number of extracted tags, frequency of appearance, importance, or a combination of at least two of these. (Appendix 4) A program described in (Appendix 2) or (Appendix 3), wherein in the step of presenting the first list, the aspect is the size, color, font, display direction of the characters of the words that make up the tag, or a combination of at least two of these. (Appendix 5) A program described in any of (Appendix 1) to (Appendix 4) that causes the processor to execute the steps of accepting a selection of one or more of the tags from the first list and presenting image-related information about the images that include the selected one or more of the tags in their description. (Appendix 6) The program according to claim 5, wherein in the step of presenting the image-related information, the image-related information includes the description. (Appendix 7) The program according to claim 6, wherein in the step of presenting the image-related information, the image-related information includes the image. (Appendix 8) A program described in any of (Appendix 1) to (Appendix 4) that causes the processor to execute the steps of accepting a selection of the tag from the first list, repeating the process of presenting the cluster to which the selected tag belongs as a second list, and presenting image-related information about the image that includes the selected tag in the description. (Appendix 9) The program according to claim 8, wherein the step of presenting the image-related information includes presenting the number of pieces of image-related information related to the second list. (Appendix 10) The program according to claim 9, wherein, in the step of accepting a selection, each time the process of accepting a selection is repeated, in the step of presenting the image-related information, the number of items to be presented is reduced (Supplementary Note 9). (Appendix 11) A program described in (Appendix 9) or (Appendix 10) that causes the processor to execute a step of accepting a request for presentation of the image-related information, and in the step of presenting the image-related information, presents the image-related information based on the accepted request. (Appendix 12) A program described in any one of (Appendix 1) to (Appendix 11), wherein in the output step, for each of the plurality of input images, text information, the image, and the instruction statement instructing the output of the explanatory text taking into account the content of the text information are input to a multimodal generative AI model, and the explanatory text is output from the multimodal generative AI model. (Appendix 13) The program according to claim 12, wherein in the outputting step, the text information includes information about a facility associated with the image. (Appendix 14) The program according to claim 12 or 13, wherein in the outputting step, the text information includes reviews by users of a facility related to the image. (Appendix 15) The program according to any one of (Supplementary Note 1) to (Supplementary Note 14), wherein in the step of receiving an input, each of the plurality of images is an image related to a restaurant. (Appendix 16) The program according to any one of (Supplementary Note 1) to (Supplementary Note 14), wherein in the step of receiving the input, each of the plurality of images is an image relating to a real estate property. (Appendix 17) A method executed by a computer having a processor and a memory, wherein the processor executes all of the steps performed in any of the inventions according to (Appendix 1) to (Appendix 16). (Appendix 18) An information processing device comprising a control unit and a storage unit, wherein the control unit executes all of the steps executed in any of the inventions according to (Appendix 1) to (Appendix 16). (Appendix 19) A system comprising means for executing all steps performed in any of the inventions according to (Appendix 1) to (Appendix 16). [Explanation of symbols]
[0113] 1. System 10...Terminal device 120…Communications Department 13...Input device 131...Touch-sensitive devices 14...Output device 15...Memory 16…Storage 19...Processor 20...Server 22...Communication IF 23...Input / output IF 25…Memory 26…Storage 29...Processor 40, 401...First tag cloud (first list) 41, 411… Tags 50...Image related information
Claims
1. A program to be executed by a computer including a processor and a memory, the program causing the processor to: accepting input of a plurality of images; For each of the plurality of input images, inputting the image and an instruction statement instructing output of an explanatory statement regarding the image into a multimodal generative AI model, and outputting the explanatory statement from the multimodal generative AI model; analyzing the outputted descriptions and extracting a plurality of characteristic words as tags; presenting the extracted plurality of tags as a first list; A program that executes the following.
2. causing the processor to analyze the extracted plurality of tags; 2. The program according to claim 1, wherein in the step of presenting the first list, the first list is presented in a manner corresponding to the analysis result in the step of analyzing.
3. 3. The program according to claim 2, wherein in the step of presenting the first list, the manner in which the first list is presented according to the analysis result is a manner in which the number of extracted tags, the frequency of appearance, the importance, or a combination of at least two of these.
4. 3. The program according to claim 2, wherein in the step of presenting the first list, the aspect is a size, color, font, or display direction of the characters of the words constituting the tag, or a combination of at least two of these.
5. accepting a selection of one or more of the tags from the first list; presenting image-related information about the image that includes the selected one or more tags in the description; The program according to claim 1, which causes the processor to execute the following steps.
6. 6. The program according to claim 5, wherein in the step of presenting the image-related information, the image-related information includes the description.
7. 7. The program according to claim 6, wherein in the step of presenting the image-related information, the image-related information includes the image.
8. accepting a selection of the tag from the first list; repeating the process of presenting the clusters to which the selected tags belong as a second list, and presenting image-related information about the images whose descriptions include the selected tags; The program according to claim 1, which causes the processor to execute the following steps.
9. 9. The program according to claim 8, wherein the step of presenting the image-related information includes presenting the number of the image-related information items related to the second list.
10. In the step of accepting the selection, each time the process of accepting the selection is repeated, 10. The program according to claim 9, wherein the number of items to be presented is reduced in the step of presenting the image-related information.
11. causing the processor to execute a step of accepting a request for presentation of the image-related information; 10. The program according to claim 9, wherein in the step of presenting the image-related information, the image-related information is presented based on the received request.
12. 2. The program according to claim 1, wherein in the output step, for each of the plurality of input images, text information, the image, and the instruction statement instructing the output of the explanatory text taking into account the content of the text information are input to a multimodal generation AI model.
13. 13. The program according to claim 12, wherein in the outputting step, the text information includes information about a facility associated with the image.
14. 13. The program according to claim 12, wherein in the outputting step, the text information includes word-of-mouth reviews by users of facilities related to the image.
15. The program according to claim 1 , wherein in the step of accepting input, each of the plurality of images is an image relating to a restaurant.
16. 2. The program according to claim 1, wherein in the step of accepting input, each of the plurality of images is an image relating to a real estate property.
17. A method implemented on a computer having a processor and a memory, wherein the processor performs all of the steps performed in the invention according to any one of claims 1 to 16.
18. 17. An information processing device comprising a control unit and a storage unit, wherein the control unit executes all of the steps executed in the invention according to any one of claims 1 to 16.
19. A system comprising means for executing all steps performed in the invention according to any one of claims 1 to 16.
Citation Information
Patent Citations
Information processing apparatus and program
JP2010160688A
Information processing apparatus, information processing method and information processing program
JP2023170790A
Technique for ranking content item recommendations - Patents.com
JP2022505237A