System
The system helps users express vague images concretely, facilitating product and information searches, customization, and commercialization requests through AI-generated images and concierge support.
Patent Information
- Application Number
- JP2024137371
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Users struggle to concretely express vague images in their minds, making it difficult to find specific products and information, and lack effective means to communicate their desires to companies when desired products are not available.
A system that allows users to input mental images in language, generates images using AI, performs image searches, and displays results, with options for customization and linking requests to companies, and includes an AI concierge for suggestions.
Enables users to easily find products and information based on their images, facilitates product customization, and communicates requests to companies for commercialization, providing appropriate advice when needed.
Smart Images

Figure 2026034250000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Users are unable to concretely express the vague image they have in their minds, making it difficult to find the products and information they want. Another issue is that if the product a user wants doesn't exist on the market, they have no way to properly communicate their desires to companies. Furthermore, if users can't solidify their image in a concrete way, they have few ways to obtain appropriate advice. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system including: a means for a user to input an image in their mind using language; a means for a server to receive the input and generate an image based on the language data using an AI model; a means for the server to perform an image search using the generated image to search for related products and information; and a means for a terminal to display the search results to the user. Furthermore, the system also includes a means for the user to input additional customization requests based on the image search results; a means for the server to receive the customization requests and generate customized images using an AI model again; a means for the server to perform another image search using the customized image; and a means for the terminal to display the customized search results to the user, making it easier for users to find specific products they desire. Furthermore, if a product desired by a user is not available, the system also includes a means for the server to link the request to companies that are interested in the product and a means for companies to consider commercializing the product based on the request, thereby enabling users to appropriately communicate their requests, including those for products not currently available on the market, to companies. Furthermore, when the user is unable to solidify an image in their mind, the device has a means for accessing an AI concierge, a means for the server to use the AI concierge to suggest an image based on the user's consultation, and a means for the device to display the suggested image to the user, making it easier for the user to obtain appropriate advice.
[0006] "User" refers to a person who uses the system.
[0007] "Terminal" refers to a device used by a user that acts as an input and output interface.
[0008] "Server" refers to a central computer that receives requests from terminals, processes data, and provides information.
[0009] "Language data" refers to text and language information entered by a user.
[0010] An "AI model" is a computational model that uses artificial intelligence to generate images and analyze data based on input data.
[0011] "Image" refers to a visual image generated by an AI model.
[0012] An "image search engine" refers to a system that searches for related information or products based on an input image.
[0013] "Search results" refers to the list of related information or products returned by an image search engine.
[0014] A "customization request" refers to an input made by a user who is not satisfied with the initial search results and specifies additional conditions or wishes.
[0015] "AI Concierge" refers to an artificial intelligence system that provides support to help users solidify their specific image.
[0016] "Enterprise" refers to a corporation or organization that provides products or services based on user requests. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] This invention is a system that allows users to input images in their minds that they cannot express concretely in words, and find appropriate information and products based on those images. The system consists of a user terminal, a server, an AI model, and an image search engine.
[0039] User Input Phase
[0040] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the device's input field. This input data is sent from the device to the server.
[0041] Image generation phase
[0042] The server receives the text data sent by the user. The server passes the received text data to the AI model, which generates an image. Specifically, the AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server.
[0043] Image search phase
[0044] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0045] Result display phase
[0046] The server sends the formatted search results to the user's device, where the user can view and check the results. This allows the user to easily find the products and information they want.
[0047] Customization Phase (Optional)
[0048] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0049] Commercialization request phase (optional)
[0050] If the product desired by the user does not exist on the market, the server will link the request to relevant companies. The server will then send the user's request and the generated image to the companies, which will then use it as a starting point for considering commercialization.
[0051] AI Concierge Phase (Optional)
[0052] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0053] As described above, the present invention allows users to concretely express the vague image in their mind, and based on that, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0057] Step 2:
[0058] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0059] Step 3:
[0060] The server receives the user's input text sent from the device via the API.
[0061] Step 4:
[0062] The server passes the received text data to the AI model and requests it to generate an image. The AI model analyzes the input text and generates a corresponding image.
[0063] Step 5:
[0064] The server receives the images generated by the AI model and stores them in a database.
[0065] Step 6:
[0066] The image stored by the server is sent to an image search engine to search for related products and information.
[0067] Step 7:
[0068] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[0069] Step 8:
[0070] The server returns an API response that sends the formatted search results to the terminal.
[0071] Step 9:
[0072] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[0073] Step 10:
[0074] If the user is not satisfied with the search results, they can enter their request for further customization into the device's input field, for example, "more brightly colored sandals."
[0075] Step 11:
[0076] The device obtains the user's customization requests and makes an API request to send them back to the server.
[0077] Step 12:
[0078] The server receives the customization request and requests the AI model to generate an image again. The AI model generates a new image based on the customized conditions.
[0079] Step 13:
[0080] The server receives the new image and causes the image search engine to search again.
[0081] Step 14:
[0082] The server receives the new search results, formats them, and sends them back to the user's device, which then displays the customized search results to the user.
[0083] Step 15:
[0084] If the product desired by the user does not exist on the market, the server will connect the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0085] Step 16:
[0086] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[0087] Step 17:
[0088] The server passes the consultation details to the AI concierge and requests a proposal. The AI concierge analyzes the consultation details and proposes the optimal image.
[0089] Step 18:
[0090] The server receives the AI concierge's suggestions and sends them to the user's device, which displays the suggested images to the user, allowing the user to make an appropriate choice.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] Conventional systems make it difficult for users to concretely express the vague image they have in their mind, which makes it difficult to find appropriate products and information. Furthermore, customization is not easy when the desired product is not available on the market or when users are dissatisfied with the search results. Furthermore, users often struggle to solidify a specific image. A system that can solve these problems is needed.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes: means for a user to input an image in their mind in language; means for the server to receive the input and generate an image based on the language data using a generative AI model; means for the server to perform an image search using the generated image to search for related products and information; means for a terminal to display the search results to the user; means for the terminal to call an AI concierge function when the user has difficulty solidifying a specific image; means for the server to pass the consultation content to the AI concierge function and propose an optimal image based on the user's wishes; and means for the terminal to display the proposal to the user. This allows users to concretely express vague images and, based on that, to search for, display, and customize information and products, and even to utilize requests for commercialization.
[0096] A "user" is someone who uses the system to input a mental image and search for products or information based on that image.
[0097] A "terminal" is a device that a user uses to input mental images as text and that displays search results and suggestions.
[0098] A "server" is a device or system that receives data sent by a user, generates images using a generative AI model, performs an image search, and returns the results to the terminal.
[0099] A "generative AI model" is an artificial intelligence model that analyzes text data entered by a user and generates visual images based on it.
[0100] A "prompt sentence" is a textual description that the user enters to concretely express the image in their mind.
[0101] An "image search engine" is a search engine that searches for related products and information based on images generated by the server.
[0102] "AI Concierge" is an artificial intelligence system that generates optimal suggestions when users are struggling with a specific choice.
[0103] A "customization request" is a user's request for further specific changes or improvements to the search results.
[0104] A "commercialization request" is a request for a product that a user desires to be produced or provided when the product does not exist on the market.
[0105] This invention is a system that allows users to concretely express the image in their mind and find appropriate information and products based on that image. This system consists of a user's device, a server, a generative AI model, and an image search engine.
[0106] User Input Phase
[0107] The user uses the device to input the image in their mind as text. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server. The device used can be a smartphone, tablet, PC, or other device, and the input field is provided in the form of a web application or mobile application.
[0108] Image generation phase
[0109] The server receives text data sent by the user. The server passes the received text data to a generative AI model (e.g., DALL-E or a similar generative model) to generate an image based on the text data. The generative AI model is a machine learning model that analyzes the text data and generates a visual image based on it. The generated image is stored on the server.
[0110] Image search phase
[0111] The generated image is sent from the server to an image search engine (for example, Google (registered trademark) image search API) to search for related products and information. The server receives the search results returned from the image search engine and formats them in a user-friendly format.
[0112] Result display phase
[0113] The server then sends the formatted search results to the user's device, which then displays the received search results for the user to review. This allows users to easily find the products and information they want.
[0114] Customization Phase
[0115] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0116] Commercialization request phase
[0117] If the product desired by the user does not exist on the market, the server will link the request to the relevant company. The server will then send the user's request and the generated image to the company, which will use it as a starting point to consider commercializing the product.
[0118] AI Concierge Phase
[0119] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to an AI concierge (e.g., GPT-4 (registered trademark)), which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0120] Specific examples
[0121] For example, if a user imagines a "retro-themed coffee shop interior," they can enter this text into their device and send it to the server. The server then uses a generative AI model to generate related images and collects information on suitable interiors through an image search engine. The results are then displayed on the device, allowing the user to view specific interior images and products.
[0122] Prompt Sentence Examples
[0123] "Think of lightweight sandals for the beach in the summer and search for related products."
[0124] "Imagine the interior of a retro coffee shop and propose a design that relates to that."
[0125] "Can you suggest an idea for a dress to wear to a colleague's wedding?"
[0126] By inputting a specific prompt sentence in this way, the system can provide appropriate information based on the user's wishes.
[0127] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0128] Step 1:
[0129] The user inputs text using the device. The user inputs the image in their mind as concrete text into the input field of the device. For example, the user might input "lightweight sandals to wear on the beach in the summer." The input data is sent from the device to the server in JSON format or HTTP request format.
[0130] Step 2:
[0131] The device sends the input data to the server as an HTTP POST request, with the text data included as the payload, which sends the image of the user's input to the server.
[0132] Step 3:
[0133] The server receives the text data. The server extracts the text data from the received HTTP request and prepares it for analysis. It receives text as input data and prepares the data to be passed to a generative AI model as output.
[0134] Step 4:
[0135] The server sends data to the generative AI model. The server creates and sends an API request to pass text data to the generative AI model (e.g., DALL-E). It sends text data as input and receives the generated image as output.
[0136] Step 5:
[0137] The AI model generates an image. The generative AI model analyzes the received text data and generates a visual image based on it. It receives text data as input and generates an image as output. The generated image is sent back to the server.
[0138] Step 6:
[0139] The server stores the generated image. The server receives the generated image and stores it in a database or file system. The image is stored in binary format for further processing.
[0140] Step 7:
[0141] The server sends the generated image to an image search engine. The server retrieves the stored image and sends it to an image search engine (e.g., Google Image Search API) in the form of a request, which sends the image as input and searches for related products and information.
[0142] Step 8:
[0143] The image search engine returns the search results. The image search engine searches for related products and information based on the received image and returns the results to the server. It takes an image as input and generates related search results as output.
[0144] Step 9:
[0145] The server formats the search results. The server converts the raw data received from the image search engine into an easy-to-understand format. It formats the product name, price, image URL, etc., and processes them into a format that can be presented to the user.
[0146] Step 10:
[0147] The server returns the formatted results to the terminal. The server then sends the formatted search results to the terminal as an HTTP response. The formatted data is sent and the results are provided to the user.
[0148] Step 11:
[0149] The device displays the search results it has received. The device analyzes the formatted search results it has received and displays them in the user interface in an appropriate layout. The user can then check products and information based on the results.
[0150] (Application example 1)
[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0152] In the past, it was difficult for users to search for products or information by associating the mental image they had in their minds, which they could not specifically describe in words. Furthermore, there was a lack of support for making vague images concrete, making it difficult for users to properly find the products they wanted. To solve this problem, a system was needed that could visually materialize the image users have in their minds and effectively search for products and information based on that image.
[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0154] In this invention, the server includes means for a user to input an image in their mind in language, means for generating an image based on the language data using a generative AI model, means for performing an image search using the generated image to search for related products and information, means for a user to input additional customization requests from a terminal, means for generating a customized image again using the generative AI model, and means for performing an image search using the customized image again. This makes it easier for a user to express a specific image, and enables products and information to be searched for and displayed based on that image.
[0155] "A means for users to input images in their minds in language" refers to an interface that allows users to input vague images or concepts in their minds as text.
[0156] "Generative AI model" refers to a machine learning model for automatically generating corresponding visual images based on specific text input.
[0157] "Image" refers to a visual representation generated from text by a generative AI model.
[0158] "Means for performing image search" refers to a function for searching for related products and information on the Internet based on the generated image.
[0159] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.
[0160] "Means for inputting customization requests" refers to an interface that allows the user to input more specific requests and conditions.
[0161] Generating a "customized image" refers to creating a new image using a generative AI model based on the customization request.
[0162] "Means for linking with companies" refers to a communication function for automatically notifying related companies of user requests and generated images.
[0163] "Commercialization methods" refers to the process by which a company analyzes or plans the manufacture and sale of a new product based on user requests.
[0164] "AI concierge function" refers to an artificial intelligence-based support function that provides optimal images and suggestions based on the user's vague requests and questions.
[0165] This invention is a system that allows users to input text images that they cannot express in concrete form, and find appropriate information and products based on those images. The system consists of a user terminal, a server, a generative AI model, and an image search engine.
[0166] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server.
[0167] The server receives text data sent by the user. The server passes the text data to a generative AI model, which generates an image. Specifically, the generative AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server. Specifically, open-source generative AI models such as DALL-E and Imagen can be applied.
[0168] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format. For this purpose, AWS (registered trademark) Rekognition or Google Cloud Vision API can be used, for example.
[0169] The server sends the formatted search results to the user's terminal, which displays the received search results so that the user can check them.
[0170] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request on their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search, and the results are provided to the user.
[0171] Furthermore, if the product desired by the user does not exist, the server can link the request to relevant companies. The server then sends the user's request and the generated image to the company, providing a starting point for the company to consider commercializing the product.
[0172] Furthermore, if a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0173] As a concrete example, consider the case where a user launches an application and enters "warm boots that won't slip on snowy winter roads." The server receives the text, and the generative AI model generates an image of "warm boots that won't slip on snowy winter roads." That image is then run through an image search engine, and related products are displayed. If the user further requests customization by adding "red boots," another customized image is generated and the results are displayed.
[0174] An example prompt is, "A user is searching for warm boots that will keep them from slipping on snowy winter roads. Please process this text input in the following format and generate an image using a generative AI model."
[0175] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0176] Step 1:
[0177] The user uses a terminal to input text representing the image in their mind. This input data is a specific description such as "Warm boots that won't slip on snowy roads in winter." This text input is sent from the user terminal to the server.
[0178] Step 2:
[0179] The server receives text data sent by the user (input). It then analyzes the text using a generative AI model and generates an image based on the input language data (data processing and data calculation). The generated image is saved on the server (output). Specifically, generative AI models such as DALL-E and Imagen are used.
[0180] Step 3:
[0181] The server sends the generated image to an image search engine (input). The image search engine searches for related products and information based on the sent image (data processing and data calculation). The search results are returned to the server (output), which formats them in a form that is easy for users to understand. Specifically, AWS Rekognition and Google Cloud Vision API are used.
[0182] Step 4:
[0183] The server sends the formatted search results to the user's terminal (input). The user's terminal receives the search results and displays them to the user (output). The user can check the displayed results.
[0184] Step 5:
[0185] If the user is not satisfied with the search results, they can input additional customization requests (input). For example, they can input a specific request such as "more brightly colored sandals." This customization request is sent from the terminal to the server.
[0186] Step 6:
[0187] The server receives the customization request (input). A new customized image is generated using the generative AI model again (data processing and data calculation). The generated image is saved on the server (output).
[0188] Step 7:
[0189] The server performs another image search using the customized image (input). The image search engine searches for related products and information based on the new image (data processing and data calculation). The search results are returned to the server (output) and reformatted again.
[0190] Step 8:
[0191] The server sends the formatted customized search results to the user's device (input). The user's device receives the search results again and displays them (output). This allows the user to check more specific products and information.
[0192] Step 9:
[0193] If the product desired by the user is not available, the user's request and the generated image are notified to related companies (input). The server then shares this information with interested companies, who then consider commercializing the product (output).
[0194] Step 10:
[0195] If the user has difficulty solidifying a specific image, they can call the AI concierge function from their device (input). The server receives the consultation details, passes them to the AI concierge, and proposes the optimal image (data processing and data calculation). The proposed image is then displayed on the user's device (output).
[0196] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0197] This invention is a system that helps users concretely express the images in their minds and provides appropriate information and products taking into account the user's emotions. This system is composed of a user terminal, a server, an AI model, an emotion engine, and an image search engine.
[0198] User Input Phase
[0199] The user uses the device to input text representing the image in their mind, for example, "lightweight sandals to wear on the beach in the summer." As the user inputs text, the device captures the input text data in real time and sends it to the server.
[0200] Emotion Recognition Phase
[0201] The server receives the text data sent from the device and passes it to the emotion engine. The emotion engine analyzes the text data and recognizes the user's emotion. For example, if the user imagines a "fun beach vacation," the emotion engine recognizes "fun."
[0202] Image generation phase
[0203] The server requests the AI model to generate an image based on the emotion data from the emotion engine. Based on the input text and emotion data, the AI model generates an image of, for example, "summer beach sandals" in bright colors that reflect "fun." The generated image is stored on the server.
[0204] Image search phase
[0205] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0206] Result display phase
[0207] The server returns an API response that sends the formatted search results to the user's device. The device displays the received search results to the user, who can then review them to find the products or information they are looking for.
[0208] Customization Phase (Optional)
[0209] If the user is not satisfied with the search results, they can input additional customization requests. For example, if the user wants "sandals in a more vibrant color," they input that request from their device and send it to the server. The server receives the customization request and generates a new image using the emotion engine and AI model. For example, it generates a new image of "sandals" that reflects "fun" in a more vibrant color. It then runs that image through an image search again and provides the results to the user.
[0210] Commercialization request phase (optional)
[0211] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0212] AI Concierge Phase (Optional)
[0213] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge and requests suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will suggest the most suitable outfit image based on that. These suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on them.
[0214] As described above, the present invention allows users to concretely express the images and emotions in their minds, and based on those images, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization. This allows users to access products and services that match their emotions and desires.
[0215] The processing flow will be explained below.
[0216] Step 1:
[0217] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0218] Step 2:
[0219] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0220] Step 3:
[0221] The server receives the user's input text sent from the device via the API.
[0222] Step 4:
[0223] The server passes the received text data to the emotion engine to recognize the user's emotion.
[0224] Step 5:
[0225] The emotion engine analyzes the text data and identifies the user's emotion, for example, recognizing the emotion that indicates "fun."
[0226] Step 6:
[0227] The server passes the recognized emotion data to the AI model and makes an image generation request.
[0228] Step 7:
[0229] Based on the input text and emotional data, the AI model generates an image of "summer beach sandals" that reflects "fun."
[0230] Step 8:
[0231] The server receives the images generated by the AI model and stores them in a database.
[0232] Step 9:
[0233] The image stored by the server is sent to an image search engine to search for related products and information.
[0234] Step 10:
[0235] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[0236] Step 11:
[0237] The server returns the formatted search results as an API response to be sent to the terminal.
[0238] Step 12:
[0239] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[0240] Step 13:
[0241] If the user is not satisfied with the search results, he or she can input additional customization requests. For example, if the user desires "sandals in brighter colors," the request is input from the terminal and transmitted to the server.
[0242] Step 14:
[0243] The device obtains the user's customization requests and makes an API request to send them back to the server.
[0244] Step 15:
[0245] The server receives the customization request and passes it back to the emotion engine to reconfirm the user's emotion.
[0246] Step 16:
[0247] The emotion engine analyzes the text data included in the customization request and recognizes the user's emotion.
[0248] Step 17:
[0249] The server again passes the recognized emotion data to the AI model and makes a customized image generation request.
[0250] Step 18:
[0251] The AI model generates new images based on customized criteria and emotional data.
[0252] Step 19:
[0253] The server receives the new image and sends it back to the image search engine.
[0254] Step 20:
[0255] The server receives the search results from the image search engine again, formats them, and sends them back to the user's terminal.
[0256] Step 21:
[0257] The terminal receives the customized search results from the server and displays them to the user.
[0258] Step 22:
[0259] If the product desired by the user is not available on the market, the terminal notifies the server of this fact.
[0260] Step 23:
[0261] The server sends the user's request and the generated image to a company, which considers commercializing the image.
[0262] Step 24:
[0263] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[0264] Step 25:
[0265] The server passes the consultation content and emotional data to the AI concierge and requests a proposal.
[0266] Step 26:
[0267] The AI concierge generates optimal suggestions based on input text and sentiment data.
[0268] Step 27:
[0269] The server receives the AI concierge's suggestions and sends them to the user's device.
[0270] Step 28:
[0271] The terminal displays the suggestions to the user, who can then make an appropriate selection.
[0272] Example 2
[0273] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0274] In the past, it was difficult for users to find a specific product that matches their mental image, and there were no systems that provided products or information that took the user's emotions into consideration. This made it difficult to provide products that matched the user's needs and emotions, and improving user satisfaction was a challenge.
[0275] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0276] In this invention, the server includes means for a user to input the image in their mind in language, means for receiving the input and analyzing the language data using an emotion recognition engine to recognize the user's emotion, means for the server to generate an image from the language data using a generative AI model based on the recognized emotion data, means for the server to use the generated image to search for related products and information using an image search engine, and means for a terminal to display the search results to the user. This makes it possible to quickly and accurately provide products and information based on the user's needs and emotions.
[0277] "User" refers to an individual or group that uses the system to materialize the images and desires in their mind.
[0278] "Terminal" refers to an input device or display device used by a user, including a personal computer, smartphone, tablet, etc.
[0279] "Server" refers to the computer system responsible for data processing, data storage, and execution of various engines and models for the entire system.
[0280] "Language data" refers to text information entered by the user, and is made up of sentences and words that express the user's mental images and wishes.
[0281] An "emotion recognition engine" refers to an algorithm or program that analyzes the language data entered by a user and identifies the user's emotions from it.
[0282] A "generative AI model" refers to an artificial intelligence model that generates output such as images or text based on input data, and performs creative generation in response to specific prompts.
[0283] A "prompt sentence" is text data input into a generative AI model, and refers to a sentence that contains instructions for generating the output image or information.
[0284] "Image" refers to a visual representation generated by a generative AI model based on the user's language and emotional data.
[0285] An "image search engine" refers to a system or program that searches the Internet for related products and information based on an input image and returns the results.
[0286] "Search results" refers to the list of related products and information returned by an image search engine, including relevant data that may be useful to the user.
[0287] "Customization request" refers to content that the user wishes to add or change based on the search results, and includes new input for regenerating or searching.
[0288] "Company" refers to a corporation or organization that considers commercialization in response to user requests.
[0289] This invention is a system that allows users to concretely express the images in their minds and provides appropriate information and products by taking into account the user's emotions. This system is composed of a user terminal, a server, an emotion recognition engine, a generative AI model, and an image search engine.
[0290] User Input Phase
[0291] The user uses the device to input the image in their mind in words. For example, they might input "lightweight sandals to wear on the beach in the summer." The device captures the text data entered by the user in real time and sends it to a server. The device can be a PC, smartphone, tablet, or other device.
[0292] Emotion Recognition Phase
[0293] The server receives the text data sent from the device and passes it to an emotion recognition engine. The emotion recognition engine (such as "EmotionAPI") analyzes the text data and recognizes the user's emotions. For example, if the user imagines a "fun beach vacation," the emotion recognition engine will identify the emotion "enjoyment."
[0294] Image generation phase
[0295] The server requests the generative AI model to generate an image based on the emotion data obtained from the emotion recognition engine. The generative AI model (e.g., "DALL-E") generates an image based on the input text and emotion data.
[0296] Examples of prompts include:
[0297] "Lightweight sandals perfect for summer beach days. Brightly colored designs for a fun beach vacation."
[0298] Based on this, the generative AI model generates a brightly colored image of "summer flip-flops," which is then stored on a server.
[0299] Image search phase
[0300] The generated image is sent from the server to an image search engine. The image search engine (for example, Google Image Search API) searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0301] Result display phase
[0302] The server sends the formatted search results to the user's device, which then displays them to the user, who can then review them to find the products or information they want.
[0303] Customization Phase (Optional)
[0304] If a user is not satisfied with the search results, they can input additional customization requests. For example, if they want "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again generates a new image using the emotion recognition engine and generative AI model. The generated image is then sent to the image search engine again, and the results are provided to the user.
[0305] Commercialization request phase (optional)
[0306] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0307] AI Concierge Phase (Optional)
[0308] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if they are unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge, which then makes optimal suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will use that information to suggest the most appropriate outfit. These suggestions are displayed on the user's device, allowing them to make an appropriate choice.
[0309] As a result, the present invention can provide products and information that are in line with the user's feelings and desires more accurately, thereby improving user satisfaction.
[0310] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0311] Step 1: User Input Phase
[0312] The user inputs the image in their mind using language. For example, they might input "lightweight sandals to wear on the beach in the summer."
[0313] The input language data is acquired in real time by the terminal, and the terminal transmits the acquired text data to the server.
[0314] Input: User text input (e.g., "Lightweight sandals for summer beach wear")
[0315] Output: Send text data from the terminal to the server
[0316] Step 2: Emotion Recognition Phase
[0317] The server passes the text data received from the device to the emotion recognition engine, which analyzes the text data and recognizes the user's emotions.
[0318] The text "Lightweight sandals to wear on the beach in summer" is passed to an emotion recognition engine, which analyzes it and extracts the emotion data "fun."
[0319] Input: Text data received by the server
[0320] Output: Emotion data extracted by the emotion recognition engine (e.g., "enjoyment")
[0321] Step 3: Image generation phase
[0322] The server sends a prompt to the generative AI model based on the emotion data obtained from the emotion recognition engine, requesting image generation. The generative AI model generates an image based on the input text and emotion data.
[0323] Prompt: "Lightweight sandals for summer beach wear. Brightly colored designs for a fun beach vacation."
[0324] Based on this prompt, the generative AI model generates a brightly colored image of "summer flip-flops" and stores the image on a server.
[0325] Input: Prompt text and emotion data sent by the server
[0326] Output: Generated image
[0327] Step 4: Image Search Phase
[0328] The generated image is sent from the server to an image search engine, which searches for related products and information and returns the search results to the server.
[0329] The server formats the received search results in a user-friendly format.
[0330] Input: Image sent by the server
[0331] Output: Related product information returned by an image search engine
[0332] Step 5: Result display phase
[0333] The server sends the formatted search results to the user's terminal, which then displays the received search results to the user.
[0334] The user checks and selects the products and information they want from the displayed search results.
[0335] Input: The formatted search results sent by the server
[0336] Output: Search results displayed on your terminal
[0337] Step 6: Customization Phase (Optional)
[0338] If the user is not satisfied with the search results, they can input additional customization requests, such as "more vibrantly colored sandals," by entering the request on their device and sending it to the server.
[0339] The server receives the customization request, generates a new image using the emotion recognition engine and generative AI model, and then sends the generated image to the image search engine again, providing the search results to the user.
[0340] Input: Customization requests entered by the user
[0341] Output: Regenerated image and new search results
[0342] Step 7: Commercialization Request Phase (Optional)
[0343] If the product desired by the user does not exist on the market, the server will link this request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0344] Input: Product development requests entered by the user
[0345] Output: Requests and images sent to the company
[0346] Step 8: AI Concierge Phase (Optional)
[0347] If the user has difficulty solidifying a specific image, they can call up the AI concierge function on their device and input their inquiry. For example, if they are unsure what to wear to a colleague's wedding, they can input their inquiry into their device.
[0348] The server acquires the consultation details and emotional data and passes them to the AI concierge. The AI concierge then makes optimal suggestions based on this information. For example, if the user is feeling "joy" or "elegance," it will suggest the perfect outfit. The suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on this information.
[0349] Input: Consultation content entered by the user
[0350] Output: Suggestions from the AI concierge
[0351] (Application example 2)
[0352] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] Conventional systems make it difficult for users to search for products based on specific product images or emotions, making it difficult to find products that fully meet the user's needs. Furthermore, with conventional text searches that are not based on emotions, there can be a gap between the products the user is looking for and the search results. The present invention aims to solve these problems by providing a more accurate product search system based on user emotions and text.
[0354] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input the image in their mind in language, means for recognizing the user's emotion from the language data using an emotion engine, means for generating an image based on the language data and emotion data using an AI model, means for the server to perform an image search using the generated image to search for related products and information, and means for the terminal to display the search results to the user. This makes it possible to quickly provide products and information based on the user's emotions and specific needs.
[0355] A "user" is an individual or entity that uses the system to accomplish a particular task.
[0356] A "mental image" is a visual representation of a particular scene or object that a user has in their mind.
[0357] "Language data" is text data that a user inputs to express an image in their mind.
[0358] A "server" is a computer system on a network that receives and processes data sent from user terminals.
[0359] An "emotion engine" is a software system that analyzes a user's language data and recognizes the emotions contained therein.
[0360] "Emotion data" is emotion information extracted from the user's language data analyzed by the emotion engine.
[0361] An "AI model" is a mathematical model that uses artificial intelligence technology to perform specific tasks.
[0362] An "image" is a visual image generated by an AI model based on a user's language and emotional data.
[0363] "Image search" is the process of searching the Internet for related products and information based on a generated image.
[0364] "Related products and information" refers to items and data found through image search based on the user's needs and emotions.
[0365] "Search results" are lists of related products and information obtained through image search.
[0366] A "terminal" is a device that can be directly operated by a user, such as a smartphone or computer.
[0367] A "customization request" is a request for improvement that a user inputs when the user is dissatisfied with the search results.
[0368] A "customized image" is an image generated again by an AI model based on customization requests.
[0369] "Interested companies" are organizations that have the potential to respond to user requests and develop or improve products.
[0370] "Marketing information" is information that includes user emotion data that companies can use as a reference when considering commercialization.
[0371] To implement this invention, the user must first input the image in their mind as concrete text. When the user inputs the text using a device such as a smartphone or computer, the data is sent to a server.
[0372] The server uses an emotion engine, such as Hugging Face's "sentiment-analysis" transformer model, to analyze the user's emotions from the received text data. The emotion engine analyzes the words and expressions in the text data and can recognize the emotions conveyed in the user's input in real time.
[0373] The server then combines the emotion data obtained from the emotion engine with the text data and generates an image using a generative AI model such as DALL-E. This is the process of generating a visual image based on the text entered by the user, further reflecting the emotion contained in the text. The generated image is then stored on the server.
[0374] The generated images are used to search for related products and information using image search engines such as the Google Image Search API. The server receives the search results returned by the image search engine and formats them in a user-friendly format.
[0375] The server sends the formatted search results to the user's device, which then displays them to the user. The user can then review the products and information displayed in the search results and make an appropriate selection.
[0376] If the user is not satisfied with the search results, they can input additional customization requests from their device. The server receives these customization requests and generates customized images using the emotion engine and generative AI model again. Further image searches are performed using these customized images, and the results are displayed to the user.
[0377] If the product desired by the user is not available on the market, the server will share the request with relevant companies. At this time, the generated image and emotion data will also be provided to the companies. Based on this, the companies can consider commercializing the product and analyze marketing information.
[0378] As a concrete example, consider a user searching for "swimsuits for a fun weekend at the beach." When the user enters this text, the emotion engine recognizes the word "fun," and the generative AI model generates a bright and cheerful image of the swimsuit. The server then searches for related products based on this image and suggests them to the user.
[0379] Example prompt sentence:
[0380] Your goal is to implement an application that recognizes emotions from the input text, such as "I'm looking for a swimsuit for a fun weekend at the beach," generates relevant images based on the emotion, and searches for and suggests related products.
[0381] In this way, the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[0382] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0383] Step 1:
[0384] The user inputs the image in their mind using language. For example, if they are looking for a swimsuit for a fun weekend at the beach, they can input this text using a device such as a smartphone or computer. The input data (text) is acquired in real time and sent to the server.
[0385] Step 2:
[0386] The server receives the text data sent from the device. Then, it uses an emotion engine (for example, Hugging Face's "sentiment-analysis" Transformer model) to analyze the user's emotions from the received text data. In this case, the input is the user's text data, and the output is emotional data such as "enjoyment."
[0387] Step 3:
[0388] The server requests a generative AI model (such as DALL-E) to generate an image based on the emotion data and text data from the emotion engine. The input is text data and emotion data, and the output is an image that reflects the emotion. Specifically, the server generates an image of a bright and cheerful beach swimsuit that reflects "fun."
[0389] Step 4:
[0390] Using the generated image, the server uses an image search engine (such as the Google Image Search API) to search for related products and information. In this case, the input is the image, and the output is a list of related products and information.
[0391] Step 5:
[0392] The server receives the search results returned by the image search engine and formats them in a user-friendly format. This formatting process ensures that the search results are displayed in a format that is appealing to the user. The input to this process is the unformatted search result data returned by the search engine, and the output is the formatted search result data.
[0393] Step 6:
[0394] The server sends the formatted search results to the user's device, which then displays the received search results to the user. The input here is the formatted search result data, and the output is the visual search results displayed on the user's device screen.
[0395] Step 7:
[0396] If the user is not satisfied with a particular product or information, they can input additional customization requests from their terminal. These requests are sent back to the server and used for the next process. The input is the text data of the user's customization requests, and the output is the customization requests sent to the server.
[0397] Step 8:
[0398] The server receives the customization request and generates a customized image using the emotion engine and generative AI model. Here, the input is the text data and emotion data of the customization request, and the output is a customized image. For example, if a user requests "sandals with a more vibrant color," the server generates an image of more vibrant sandals.
[0399] Step 9:
[0400] The server then performs another image search using the customized image, where the input is the customized image and the output is a new list of related products and information, formats the search results, and presents them to the user again.
[0401] Step 10:
[0402] If the product desired by the user does not exist on the market, the server will share this request with relevant companies. The companies will then consider commercializing the product. The input here is the user's request, the generated image, and emotion data, and the output is marketing information provided to the company and the results of consideration for commercialization.
[0403] In this way, the system of the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[0404] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0405] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0406] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0407] [Second embodiment]
[0408] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0409] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0410] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0411] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0412] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0414] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0415] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0416] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0417] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0418] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0419] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0420] This invention is a system that allows users to input images in their minds that they cannot express concretely in words, and find appropriate information and products based on those images. The system consists of a user terminal, a server, an AI model, and an image search engine.
[0421] User Input Phase
[0422] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the device's input field. This input data is sent from the device to the server.
[0423] Image generation phase
[0424] The server receives the text data sent by the user. The server passes the received text data to the AI model, which generates an image. Specifically, the AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server.
[0425] Image search phase
[0426] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0427] Result display phase
[0428] The server sends the formatted search results to the user's device, where the user can view and check the results. This allows the user to easily find the products and information they want.
[0429] Customization Phase (Optional)
[0430] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0431] Commercialization request phase (optional)
[0432] If the product desired by the user does not exist on the market, the server will link the request to relevant companies. The server will then send the user's request and the generated image to the companies, which will then use it as a starting point for considering commercialization.
[0433] AI Concierge Phase (Optional)
[0434] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0435] As described above, the present invention allows users to concretely express the vague image in their mind, and based on that, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization.
[0436] The processing flow will be explained below.
[0437] Step 1:
[0438] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0439] Step 2:
[0440] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0441] Step 3:
[0442] The server receives the user's input text sent from the device via the API.
[0443] Step 4:
[0444] The server passes the received text data to the AI model and requests it to generate an image. The AI model analyzes the input text and generates a corresponding image.
[0445] Step 5:
[0446] The server receives the images generated by the AI model and stores them in a database.
[0447] Step 6:
[0448] The image stored by the server is sent to an image search engine to search for related products and information.
[0449] Step 7:
[0450] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[0451] Step 8:
[0452] The server returns an API response that sends the formatted search results to the terminal.
[0453] Step 9:
[0454] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[0455] Step 10:
[0456] If the user is not satisfied with the search results, they can enter their request for further customization into the device's input field, for example, "more brightly colored sandals."
[0457] Step 11:
[0458] The device obtains the user's customization requests and makes an API request to send them back to the server.
[0459] Step 12:
[0460] The server receives the customization request and requests the AI model to generate an image again. The AI model generates a new image based on the customized conditions.
[0461] Step 13:
[0462] The server receives the new image and causes the image search engine to search again.
[0463] Step 14:
[0464] The server receives the new search results, formats them, and sends them back to the user's device, which then displays the customized search results to the user.
[0465] Step 15:
[0466] If the product desired by the user does not exist on the market, the server will connect the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0467] Step 16:
[0468] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[0469] Step 17:
[0470] The server passes the consultation details to the AI concierge and requests a proposal. The AI concierge analyzes the consultation details and proposes the optimal image.
[0471] Step 18:
[0472] The server receives the AI concierge's suggestions and sends them to the user's device, which displays the suggested images to the user, allowing the user to make an appropriate choice.
[0473] Example 1
[0474] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0475] Conventional systems make it difficult for users to concretely express the vague image they have in their mind, which makes it difficult to find appropriate products and information. Furthermore, customization is not easy when the desired product is not available on the market or when users are dissatisfied with the search results. Furthermore, users often struggle to solidify a specific image. A system that can solve these problems is needed.
[0476] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0477] In this invention, the server includes: means for a user to input an image in their mind in language; means for the server to receive the input and generate an image based on the language data using a generative AI model; means for the server to perform an image search using the generated image to search for related products and information; means for a terminal to display the search results to the user; means for the terminal to call an AI concierge function when the user has difficulty solidifying a specific image; means for the server to pass the consultation content to the AI concierge function and propose an optimal image based on the user's wishes; and means for the terminal to display the proposal to the user. This allows users to concretely express vague images and, based on that, to search for, display, and customize information and products, and even to utilize requests for commercialization.
[0478] A "user" is someone who uses the system to input a mental image and search for products or information based on that image.
[0479] A "terminal" is a device that a user uses to input mental images as text and that displays search results and suggestions.
[0480] A "server" is a device or system that receives data sent by a user, generates images using a generative AI model, performs an image search, and returns the results to the terminal.
[0481] A "generative AI model" is an artificial intelligence model that analyzes text data entered by a user and generates visual images based on it.
[0482] A "prompt sentence" is a textual description that the user enters to concretely express the image in their mind.
[0483] An "image search engine" is a search engine that searches for related products and information based on images generated by the server.
[0484] "AI Concierge" is an artificial intelligence system that generates optimal suggestions when users are struggling with a specific choice.
[0485] A "customization request" is a user's request for further specific changes or improvements to the search results.
[0486] A "commercialization request" is a request for a product that a user desires to be produced or provided when the product does not exist on the market.
[0487] This invention is a system that allows users to concretely express the image in their mind and find appropriate information and products based on that image. This system consists of a user's device, a server, a generative AI model, and an image search engine.
[0488] User Input Phase
[0489] The user uses the device to input the image in their mind as text. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server. The device used can be a smartphone, tablet, PC, or other device, and the input field is provided in the form of a web application or mobile application.
[0490] Image generation phase
[0491] The server receives text data sent by the user. The server passes the received text data to a generative AI model (e.g., DALL-E or a similar generative model) to generate an image based on the text data. The generative AI model is a machine learning model that analyzes the text data and generates a visual image based on it. The generated image is stored on the server.
[0492] Image search phase
[0493] The generated image is sent from the server to an image search engine (for example, Google Image Search API) to search for related products and information. The server receives the search results returned from the image search engine and formats them in a user-friendly format.
[0494] Result display phase
[0495] The server then sends the formatted search results to the user's device, which then displays the received search results for the user to review. This allows users to easily find the products and information they want.
[0496] Customization Phase
[0497] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0498] Commercialization request phase
[0499] If the product desired by the user does not exist on the market, the server will link the request to the relevant company. The server will then send the user's request and the generated image to the company, which will use it as a starting point to consider commercializing the product.
[0500] AI Concierge Phase
[0501] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to an AI concierge (e.g., GPT-4), which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0502] Specific examples
[0503] For example, if a user imagines a "retro-themed coffee shop interior," they can enter this text into their device and send it to the server. The server then uses a generative AI model to generate related images and collects information on suitable interiors through an image search engine. The results are then displayed on the device, allowing the user to view specific interior images and products.
[0504] Prompt Sentence Examples
[0505] "Think of lightweight sandals for the beach in the summer and search for related products."
[0506] "Imagine the interior of a retro coffee shop and propose a design that relates to that."
[0507] "Can you suggest an idea for a dress to wear to a colleague's wedding?"
[0508] By inputting a specific prompt sentence in this way, the system can provide appropriate information based on the user's wishes.
[0509] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0510] Step 1:
[0511] The user inputs text using the device. The user inputs the image in their mind as concrete text into the input field of the device. For example, the user might input "lightweight sandals to wear on the beach in the summer." The input data is sent from the device to the server in JSON format or HTTP request format.
[0512] Step 2:
[0513] The device sends the input data to the server as an HTTP POST request, with the text data included as the payload, which sends the image of the user's input to the server.
[0514] Step 3:
[0515] The server receives the text data. The server extracts the text data from the received HTTP request and prepares it for analysis. It receives text as input data and prepares the data to be passed to a generative AI model as output.
[0516] Step 4:
[0517] The server sends data to the generative AI model. The server creates and sends an API request to pass text data to the generative AI model (e.g., DALL-E). It sends text data as input and receives the generated image as output.
[0518] Step 5:
[0519] The AI model generates an image. The generative AI model analyzes the received text data and generates a visual image based on it. It receives text data as input and generates an image as output. The generated image is sent back to the server.
[0520] Step 6:
[0521] The server stores the generated image. The server receives the generated image and stores it in a database or file system. The image is stored in binary format for further processing.
[0522] Step 7:
[0523] The server sends the generated image to an image search engine. The server retrieves the stored image and sends it to an image search engine (e.g., Google Image Search API) in the form of a request, which sends the image as input and searches for related products and information.
[0524] Step 8:
[0525] The image search engine returns the search results. The image search engine searches for related products and information based on the received image and returns the results to the server. It takes an image as input and generates related search results as output.
[0526] Step 9:
[0527] The server formats the search results. The server converts the raw data received from the image search engine into an easy-to-understand format. It formats the product name, price, image URL, etc., and processes them into a format that can be presented to the user.
[0528] Step 10:
[0529] The server returns the formatted results to the terminal. The server then sends the formatted search results to the terminal as an HTTP response. The formatted data is sent and the results are provided to the user.
[0530] Step 11:
[0531] The device displays the search results it has received. The device analyzes the formatted search results it has received and displays them in the user interface in an appropriate layout. The user can then check products and information based on the results.
[0532] (Application example 1)
[0533] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0534] In the past, it was difficult for users to search for products or information by associating the mental image they had in their minds, which they could not specifically describe in words. Furthermore, there was a lack of support for making vague images concrete, making it difficult for users to properly find the products they wanted. To solve this problem, a system was needed that could visually materialize the image users have in their minds and effectively search for products and information based on that image.
[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0536] In this invention, the server includes means for a user to input an image in their mind in language, means for generating an image based on the language data using a generative AI model, means for performing an image search using the generated image to search for related products and information, means for a user to input additional customization requests from a terminal, means for generating a customized image again using the generative AI model, and means for performing an image search using the customized image again. This makes it easier for a user to express a specific image, and enables products and information to be searched for and displayed based on that image.
[0537] "A means for users to input images in their minds in language" refers to an interface that allows users to input vague images or concepts in their minds as text.
[0538] "Generative AI model" refers to a machine learning model for automatically generating corresponding visual images based on specific text input.
[0539] "Image" refers to a visual representation generated from text by a generative AI model.
[0540] "Means for performing image search" refers to a function for searching for related products and information on the Internet based on the generated image.
[0541] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.
[0542] "Means for inputting customization requests" refers to an interface that allows the user to input more specific requests and conditions.
[0543] Generating a "customized image" refers to creating a new image using a generative AI model based on the customization request.
[0544] "Means for linking with companies" refers to a communication function for automatically notifying related companies of user requests and generated images.
[0545] "Commercialization methods" refers to the process by which a company analyzes or plans the manufacture and sale of a new product based on user requests.
[0546] "AI concierge function" refers to an artificial intelligence-based support function that provides optimal images and suggestions based on the user's vague requests and questions.
[0547] This invention is a system that allows users to input text images that they cannot express in concrete form, and find appropriate information and products based on those images. The system consists of a user terminal, a server, a generative AI model, and an image search engine.
[0548] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server.
[0549] The server receives text data sent by the user. The server passes the text data to a generative AI model, which generates an image. Specifically, the generative AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server. Specifically, open-source generative AI models such as DALL-E and Imagen can be applied.
[0550] The generated image is sent from the server to an image search engine, which searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format. For this purpose, AWS Rekognition or Google Cloud Vision API can be used, for example.
[0551] The server sends the formatted search results to the user's terminal, which displays the received search results so that the user can check them.
[0552] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request on their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search, and the results are provided to the user.
[0553] Furthermore, if the product desired by the user does not exist, the server can link the request to relevant companies. The server then sends the user's request and the generated image to the company, providing a starting point for the company to consider commercializing the product.
[0554] Furthermore, if a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0555] As a concrete example, consider the case where a user launches an application and enters "warm boots that won't slip on snowy winter roads." The server receives the text, and the generative AI model generates an image of "warm boots that won't slip on snowy winter roads." That image is then run through an image search engine, and related products are displayed. If the user further requests customization by adding "red boots," another customized image is generated and the results are displayed.
[0556] An example prompt is, "A user is searching for warm boots that will keep them from slipping on snowy winter roads. Please process this text input in the following format and generate an image using a generative AI model."
[0557] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0558] Step 1:
[0559] The user uses a terminal to input text representing the image in their mind. This input data is a specific description such as "Warm boots that won't slip on snowy roads in winter." This text input is sent from the user terminal to the server.
[0560] Step 2:
[0561] The server receives text data sent by the user (input). It then analyzes the text using a generative AI model and generates an image based on the input language data (data processing and data calculation). The generated image is saved on the server (output). Specifically, generative AI models such as DALL-E and Imagen are used.
[0562] Step 3:
[0563] The server sends the generated image to an image search engine (input). The image search engine searches for related products and information based on the sent image (data processing and data calculation). The search results are returned to the server (output), which formats them in a form that is easy for users to understand. Specifically, AWS Rekognition and Google Cloud Vision API are used.
[0564] Step 4:
[0565] The server sends the formatted search results to the user's terminal (input). The user's terminal receives the search results and displays them to the user (output). The user can check the displayed results.
[0566] Step 5:
[0567] If the user is not satisfied with the search results, they can input additional customization requests (input). For example, they can input a specific request such as "more brightly colored sandals." This customization request is sent from the terminal to the server.
[0568] Step 6:
[0569] The server receives the customization request (input). A new customized image is generated using the generative AI model again (data processing and data calculation). The generated image is saved on the server (output).
[0570] Step 7:
[0571] The server performs another image search using the customized image (input). The image search engine searches for related products and information based on the new image (data processing and data calculation). The search results are returned to the server (output) and reformatted again.
[0572] Step 8:
[0573] The server sends the formatted customized search results to the user's device (input). The user's device receives the search results again and displays them (output). This allows the user to check more specific products and information.
[0574] Step 9:
[0575] If the product desired by the user is not available, the user's request and the generated image are notified to related companies (input). The server then shares this information with interested companies, who then consider commercializing the product (output).
[0576] Step 10:
[0577] If the user has difficulty solidifying a specific image, they can call the AI concierge function from their device (input). The server receives the consultation details, passes them to the AI concierge, and proposes the optimal image (data processing and data calculation). The proposed image is then displayed on the user's device (output).
[0578] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0579] This invention is a system that helps users concretely express the images in their minds and provides appropriate information and products taking into account the user's emotions. This system is composed of a user terminal, a server, an AI model, an emotion engine, and an image search engine.
[0580] User Input Phase
[0581] The user uses the device to input text representing the image in their mind, for example, "lightweight sandals to wear on the beach in the summer." As the user inputs text, the device captures the input text data in real time and sends it to the server.
[0582] Emotion Recognition Phase
[0583] The server receives the text data sent from the device and passes it to the emotion engine. The emotion engine analyzes the text data and recognizes the user's emotion. For example, if the user imagines a "fun beach vacation," the emotion engine recognizes "fun."
[0584] Image generation phase
[0585] The server requests the AI model to generate an image based on the emotion data from the emotion engine. Based on the input text and emotion data, the AI model generates an image of, for example, "summer beach sandals" in bright colors that reflect "fun." The generated image is stored on the server.
[0586] Image search phase
[0587] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0588] Result display phase
[0589] The server returns an API response that sends the formatted search results to the user's device. The device displays the received search results to the user, who can then review them to find the products or information they are looking for.
[0590] Customization Phase (Optional)
[0591] If the user is not satisfied with the search results, they can input additional customization requests. For example, if the user wants "sandals in a more vibrant color," they input that request from their device and send it to the server. The server receives the customization request and generates a new image using the emotion engine and AI model. For example, it generates a new image of "sandals" that reflects "fun" in a more vibrant color. It then runs that image through an image search again and provides the results to the user.
[0592] Commercialization request phase (optional)
[0593] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0594] AI Concierge Phase (Optional)
[0595] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge and requests suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will suggest the most suitable outfit image based on that. These suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on them.
[0596] As described above, the present invention allows users to concretely express the images and emotions in their minds, and based on those images, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization. This allows users to access products and services that match their emotions and desires.
[0597] The processing flow will be explained below.
[0598] Step 1:
[0599] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0600] Step 2:
[0601] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0602] Step 3:
[0603] The server receives the user's input text sent from the device via the API.
[0604] Step 4:
[0605] The server passes the received text data to the emotion engine to recognize the user's emotion.
[0606] Step 5:
[0607] The emotion engine analyzes the text data and identifies the user's emotion, for example, recognizing the emotion that indicates "fun."
[0608] Step 6:
[0609] The server passes the recognized emotion data to the AI model and makes an image generation request.
[0610] Step 7:
[0611] Based on the input text and emotional data, the AI model generates an image of "summer beach sandals" that reflects "fun."
[0612] Step 8:
[0613] The server receives the images generated by the AI model and stores them in a database.
[0614] Step 9:
[0615] The image stored by the server is sent to an image search engine to search for related products and information.
[0616] Step 10:
[0617] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[0618] Step 11:
[0619] The server returns the formatted search results as an API response to be sent to the terminal.
[0620] Step 12:
[0621] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[0622] Step 13:
[0623] If the user is not satisfied with the search results, he or she can input additional customization requests. For example, if the user desires "sandals in brighter colors," the request is input from the terminal and transmitted to the server.
[0624] Step 14:
[0625] The device obtains the user's customization requests and makes an API request to send them back to the server.
[0626] Step 15:
[0627] The server receives the customization request and passes it back to the emotion engine to reconfirm the user's emotion.
[0628] Step 16:
[0629] The emotion engine analyzes the text data included in the customization request and recognizes the user's emotion.
[0630] Step 17:
[0631] The server again passes the recognized emotion data to the AI model and makes a customized image generation request.
[0632] Step 18:
[0633] The AI model generates new images based on customized criteria and emotional data.
[0634] Step 19:
[0635] The server receives the new image and sends it back to the image search engine.
[0636] Step 20:
[0637] The server receives the search results from the image search engine again, formats them, and sends them back to the user's terminal.
[0638] Step 21:
[0639] The terminal receives the customized search results from the server and displays them to the user.
[0640] Step 22:
[0641] If the product desired by the user is not available on the market, the terminal notifies the server of this fact.
[0642] Step 23:
[0643] The server sends the user's request and the generated image to a company, which considers commercializing the image.
[0644] Step 24:
[0645] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[0646] Step 25:
[0647] The server passes the consultation content and emotional data to the AI concierge and requests a proposal.
[0648] Step 26:
[0649] The AI concierge generates optimal suggestions based on input text and sentiment data.
[0650] Step 27:
[0651] The server receives the AI concierge's suggestions and sends them to the user's device.
[0652] Step 28:
[0653] The terminal displays the suggestions to the user, who can then make an appropriate selection.
[0654] Example 2
[0655] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0656] In the past, it was difficult for users to find a specific product that matches their mental image, and there were no systems that provided products or information that took the user's emotions into consideration. This made it difficult to provide products that matched the user's needs and emotions, and improving user satisfaction was a challenge.
[0657] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0658] In this invention, the server includes means for a user to input the image in their mind in language, means for receiving the input and analyzing the language data using an emotion recognition engine to recognize the user's emotion, means for the server to generate an image from the language data using a generative AI model based on the recognized emotion data, means for the server to use the generated image to search for related products and information using an image search engine, and means for a terminal to display the search results to the user. This makes it possible to quickly and accurately provide products and information based on the user's needs and emotions.
[0659] "User" refers to an individual or group that uses the system to materialize the images and desires in their mind.
[0660] "Terminal" refers to an input device or display device used by a user, including a personal computer, smartphone, tablet, etc.
[0661] "Server" refers to the computer system responsible for data processing, data storage, and execution of various engines and models for the entire system.
[0662] "Language data" refers to text information entered by the user, and is made up of sentences and words that express the user's mental images and wishes.
[0663] An "emotion recognition engine" refers to an algorithm or program that analyzes the language data entered by a user and identifies the user's emotions from it.
[0664] A "generative AI model" refers to an artificial intelligence model that generates output such as images or text based on input data, and performs creative generation in response to specific prompts.
[0665] A "prompt sentence" is text data input into a generative AI model, and refers to a sentence that contains instructions for generating the output image or information.
[0666] "Image" refers to a visual representation generated by a generative AI model based on the user's language and emotional data.
[0667] An "image search engine" refers to a system or program that searches the Internet for related products and information based on an input image and returns the results.
[0668] "Search results" refers to the list of related products and information returned by an image search engine, including relevant data that may be useful to the user.
[0669] "Customization request" refers to content that the user wishes to add or change based on the search results, and includes new input for regenerating or searching.
[0670] "Company" refers to a corporation or organization that considers commercialization in response to user requests.
[0671] This invention is a system that allows users to concretely express the images in their minds and provides appropriate information and products by taking into account the user's emotions. This system is composed of a user terminal, a server, an emotion recognition engine, a generative AI model, and an image search engine.
[0672] User Input Phase
[0673] The user uses the device to input the image in their mind in words. For example, they might input "lightweight sandals to wear on the beach in the summer." The device captures the text data entered by the user in real time and sends it to a server. The device can be a PC, smartphone, tablet, or other device.
[0674] Emotion Recognition Phase
[0675] The server receives the text data sent from the device and passes it to an emotion recognition engine. The emotion recognition engine (such as "EmotionAPI") analyzes the text data and recognizes the user's emotions. For example, if the user imagines a "fun beach vacation," the emotion recognition engine will identify the emotion "enjoyment."
[0676] Image generation phase
[0677] The server requests the generative AI model to generate an image based on the emotion data obtained from the emotion recognition engine. The generative AI model (e.g., "DALL-E") generates an image based on the input text and emotion data.
[0678] Examples of prompts include:
[0679] "Lightweight sandals perfect for summer beach days. Brightly colored designs for a fun beach vacation."
[0680] Based on this, the generative AI model generates a brightly colored image of "summer flip-flops," which is then stored on a server.
[0681] Image search phase
[0682] The generated image is sent from the server to an image search engine. The image search engine (for example, Google Image Search API) searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0683] Result display phase
[0684] The server sends the formatted search results to the user's device, which then displays them to the user, who can then review them to find the products or information they want.
[0685] Customization Phase (Optional)
[0686] If a user is not satisfied with the search results, they can input additional customization requests. For example, if they want "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again generates a new image using the emotion recognition engine and generative AI model. The generated image is then sent to the image search engine again, and the results are provided to the user.
[0687] Commercialization request phase (optional)
[0688] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0689] AI Concierge Phase (Optional)
[0690] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if they are unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge, which then makes optimal suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will use that information to suggest the most appropriate outfit. These suggestions are displayed on the user's device, allowing them to make an appropriate choice.
[0691] As a result, the present invention can provide products and information that are in line with the user's feelings and desires more accurately, thereby improving user satisfaction.
[0692] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0693] Step 1: User Input Phase
[0694] The user inputs the image in their mind using language. For example, they might input "lightweight sandals to wear on the beach in the summer."
[0695] The input language data is acquired in real time by the terminal, and the terminal transmits the acquired text data to the server.
[0696] Input: User text input (e.g., "Lightweight sandals for summer beach wear")
[0697] Output: Send text data from the terminal to the server
[0698] Step 2: Emotion Recognition Phase
[0699] The server passes the text data received from the device to the emotion recognition engine, which analyzes the text data and recognizes the user's emotions.
[0700] The text "Lightweight sandals to wear on the beach in summer" is passed to an emotion recognition engine, which analyzes it and extracts the emotion data "fun."
[0701] Input: Text data received by the server
[0702] Output: Emotion data extracted by the emotion recognition engine (e.g., "enjoyment")
[0703] Step 3: Image generation phase
[0704] The server sends a prompt to the generative AI model based on the emotion data obtained from the emotion recognition engine, requesting image generation. The generative AI model generates an image based on the input text and emotion data.
[0705] Prompt: "Lightweight sandals for summer beach wear. Brightly colored designs for a fun beach vacation."
[0706] Based on this prompt, the generative AI model generates a brightly colored image of "summer flip-flops" and stores the image on a server.
[0707] Input: Prompt text and emotion data sent by the server
[0708] Output: Generated image
[0709] Step 4: Image Search Phase
[0710] The generated image is sent from the server to an image search engine, which searches for related products and information and returns the search results to the server.
[0711] The server formats the received search results in a user-friendly format.
[0712] Input: Image sent by the server
[0713] Output: Related product information returned by an image search engine
[0714] Step 5: Result display phase
[0715] The server sends the formatted search results to the user's terminal, which then displays the received search results to the user.
[0716] The user checks and selects the products and information they want from the displayed search results.
[0717] Input: The formatted search results sent by the server
[0718] Output: Search results displayed on your terminal
[0719] Step 6: Customization Phase (Optional)
[0720] If the user is not satisfied with the search results, they can input additional customization requests, such as "more vibrantly colored sandals," by entering the request on their device and sending it to the server.
[0721] The server receives the customization request, generates a new image using the emotion recognition engine and generative AI model, and then sends the generated image to the image search engine again, providing the search results to the user.
[0722] Input: Customization requests entered by the user
[0723] Output: Regenerated image and new search results
[0724] Step 7: Commercialization Request Phase (Optional)
[0725] If the product desired by the user does not exist on the market, the server will link this request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0726] Input: Product development requests entered by the user
[0727] Output: Requests and images sent to the company
[0728] Step 8: AI Concierge Phase (Optional)
[0729] If the user has difficulty solidifying a specific image, they can call up the AI concierge function on their device and input their inquiry. For example, if they are unsure what to wear to a colleague's wedding, they can input their inquiry into their device.
[0730] The server acquires the consultation details and emotional data and passes them to the AI concierge. The AI concierge then makes optimal suggestions based on this information. For example, if the user is feeling "joy" or "elegance," it will suggest the perfect outfit. The suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on this information.
[0731] Input: Consultation content entered by the user
[0732] Output: Suggestions from the AI concierge
[0733] (Application example 2)
[0734] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0735] Conventional systems make it difficult for users to search for products based on specific product images or emotions, making it difficult to find products that fully meet the user's needs. Furthermore, with conventional text searches that are not based on emotions, there can be a gap between the products the user is looking for and the search results. The present invention aims to solve these problems by providing a more accurate product search system based on user emotions and text.
[0736] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input the image in their mind in language, means for recognizing the user's emotion from the language data using an emotion engine, means for generating an image based on the language data and emotion data using an AI model, means for the server to perform an image search using the generated image to search for related products and information, and means for the terminal to display the search results to the user. This makes it possible to quickly provide products and information based on the user's emotions and specific needs.
[0737] A "user" is an individual or entity that uses the system to accomplish a particular task.
[0738] A "mental image" is a visual representation of a particular scene or object that a user has in their mind.
[0739] "Language data" is text data that a user inputs to express an image in their mind.
[0740] A "server" is a computer system on a network that receives and processes data sent from user terminals.
[0741] An "emotion engine" is a software system that analyzes a user's language data and recognizes the emotions contained therein.
[0742] "Emotion data" is emotion information extracted from the user's language data analyzed by the emotion engine.
[0743] An "AI model" is a mathematical model that uses artificial intelligence technology to perform specific tasks.
[0744] An "image" is a visual image generated by an AI model based on a user's language and emotional data.
[0745] "Image search" is the process of searching the Internet for related products and information based on a generated image.
[0746] "Related products and information" refers to items and data found through image search based on the user's needs and emotions.
[0747] "Search results" are lists of related products and information obtained through image search.
[0748] A "terminal" is a device that can be directly operated by a user, such as a smartphone or computer.
[0749] A "customization request" is a request for improvement that a user inputs when the user is dissatisfied with the search results.
[0750] A "customized image" is an image generated again by an AI model based on customization requests.
[0751] "Interested companies" are organizations that have the potential to respond to user requests and develop or improve products.
[0752] "Marketing information" is information that includes user emotion data that companies can use as a reference when considering commercialization.
[0753] To implement this invention, the user must first input the image in their mind as concrete text. When the user inputs the text using a device such as a smartphone or computer, the data is sent to a server.
[0754] The server uses an emotion engine, such as Hugging Face's "sentiment-analysis" transformer model, to analyze the user's emotions from the received text data. The emotion engine analyzes the words and expressions in the text data and can recognize the emotions conveyed in the user's input in real time.
[0755] The server then combines the emotion data obtained from the emotion engine with the text data and generates an image using a generative AI model such as DALL-E. This is the process of generating a visual image based on the text entered by the user, further reflecting the emotion contained in the text. The generated image is then stored on the server.
[0756] The generated images are used to search for related products and information using image search engines such as the Google Image Search API. The server receives the search results returned by the image search engine and formats them in a user-friendly format.
[0757] The server sends the formatted search results to the user's device, which then displays them to the user. The user can then review the products and information displayed in the search results and make an appropriate selection.
[0758] If the user is not satisfied with the search results, they can input additional customization requests from their device. The server receives these customization requests and generates customized images using the emotion engine and generative AI model again. Further image searches are performed using these customized images, and the results are displayed to the user.
[0759] If the product desired by the user is not available on the market, the server will share the request with relevant companies. At this time, the generated image and emotion data will also be provided to the companies. Based on this, the companies can consider commercializing the product and analyze marketing information.
[0760] As a concrete example, consider a user searching for "swimsuits for a fun weekend at the beach." When the user enters this text, the emotion engine recognizes the word "fun," and the generative AI model generates a bright and cheerful image of the swimsuit. The server then searches for related products based on this image and suggests them to the user.
[0761] Example prompt sentence:
[0762] Your goal is to implement an application that recognizes emotions from the input text, such as "I'm looking for a swimsuit for a fun weekend at the beach," generates relevant images based on the emotion, and searches for and suggests related products.
[0763] In this way, the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[0764] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0765] Step 1:
[0766] The user inputs the image in their mind using language. For example, if they are looking for a swimsuit for a fun weekend at the beach, they can input this text using a device such as a smartphone or computer. The input data (text) is acquired in real time and sent to the server.
[0767] Step 2:
[0768] The server receives the text data sent from the device. Then, it uses an emotion engine (for example, Hugging Face's "sentiment-analysis" Transformer model) to analyze the user's emotions from the received text data. In this case, the input is the user's text data, and the output is emotional data such as "enjoyment."
[0769] Step 3:
[0770] The server requests a generative AI model (such as DALL-E) to generate an image based on the emotion data and text data from the emotion engine. The input is text data and emotion data, and the output is an image that reflects the emotion. Specifically, the server generates an image of a bright and cheerful beach swimsuit that reflects "fun."
[0771] Step 4:
[0772] Using the generated image, the server uses an image search engine (such as the Google Image Search API) to search for related products and information. In this case, the input is the image, and the output is a list of related products and information.
[0773] Step 5:
[0774] The server receives the search results returned by the image search engine and formats them in a user-friendly format. This formatting process ensures that the search results are displayed in a format that is appealing to the user. The input to this process is the unformatted search result data returned by the search engine, and the output is the formatted search result data.
[0775] Step 6:
[0776] The server sends the formatted search results to the user's device, which then displays the received search results to the user. The input here is the formatted search result data, and the output is the visual search results displayed on the user's device screen.
[0777] Step 7:
[0778] If the user is not satisfied with a particular product or information, they can input additional customization requests from their terminal. These requests are sent back to the server and used for the next process. The input is the text data of the user's customization requests, and the output is the customization requests sent to the server.
[0779] Step 8:
[0780] The server receives the customization request and generates a customized image using the emotion engine and generative AI model. Here, the input is the text data and emotion data of the customization request, and the output is a customized image. For example, if a user requests "sandals with a more vibrant color," the server generates an image of more vibrant sandals.
[0781] Step 9:
[0782] The server then performs another image search using the customized image, where the input is the customized image and the output is a new list of related products and information, formats the search results, and presents them to the user again.
[0783] Step 10:
[0784] If the product desired by the user does not exist on the market, the server will share this request with relevant companies. The companies will then consider commercializing the product. The input here is the user's request, the generated image, and emotion data, and the output is marketing information provided to the company and the results of consideration for commercialization.
[0785] In this way, the system of the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[0786] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0787] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0788] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0789] [Third embodiment]
[0790] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0791] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0792] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0793] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0794] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0795] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0796] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0797] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0798] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0799] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0800] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0801] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0802] This invention is a system that allows users to input images in their minds that they cannot express concretely in words, and find appropriate information and products based on those images. The system consists of a user terminal, a server, an AI model, and an image search engine.
[0803] User Input Phase
[0804] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the device's input field. This input data is sent from the device to the server.
[0805] Image generation phase
[0806] The server receives the text data sent by the user. The server passes the received text data to the AI model, which generates an image. Specifically, the AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server.
[0807] Image search phase
[0808] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0809] Result display phase
[0810] The server sends the formatted search results to the user's device, where the user can view and check the results. This allows the user to easily find the products and information they want.
[0811] Customization Phase (Optional)
[0812] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0813] Commercialization request phase (optional)
[0814] If the product desired by the user does not exist on the market, the server will link the request to relevant companies. The server will then send the user's request and the generated image to the companies, which will then use it as a starting point for considering commercialization.
[0815] AI Concierge Phase (Optional)
[0816] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0817] As described above, the present invention allows users to concretely express the vague image in their mind, and based on that, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization.
[0818] The processing flow will be explained below.
[0819] Step 1:
[0820] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0821] Step 2:
[0822] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0823] Step 3:
[0824] The server receives the user's input text sent from the device via the API.
[0825] Step 4:
[0826] The server passes the received text data to the AI model and requests it to generate an image. The AI model analyzes the input text and generates a corresponding image.
[0827] Step 5:
[0828] The server receives the images generated by the AI model and stores them in a database.
[0829] Step 6:
[0830] The image stored by the server is sent to an image search engine to search for related products and information.
[0831] Step 7:
[0832] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[0833] Step 8:
[0834] The server returns an API response that sends the formatted search results to the terminal.
[0835] Step 9:
[0836] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[0837] Step 10:
[0838] If the user is not satisfied with the search results, they can enter their request for further customization into the device's input field, for example, "more brightly colored sandals."
[0839] Step 11:
[0840] The device obtains the user's customization requests and makes an API request to send them back to the server.
[0841] Step 12:
[0842] The server receives the customization request and requests the AI model to generate an image again. The AI model generates a new image based on the customized conditions.
[0843] Step 13:
[0844] The server receives the new image and causes the image search engine to search again.
[0845] Step 14:
[0846] The server receives the new search results, formats them, and sends them back to the user's device, which then displays the customized search results to the user.
[0847] Step 15:
[0848] If the product desired by the user does not exist on the market, the server will connect the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0849] Step 16:
[0850] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[0851] Step 17:
[0852] The server passes the consultation details to the AI concierge and requests a proposal. The AI concierge analyzes the consultation details and proposes the optimal image.
[0853] Step 18:
[0854] The server receives the AI concierge's suggestions and sends them to the user's device, which displays the suggested images to the user, allowing the user to make an appropriate choice.
[0855] Example 1
[0856] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0857] Conventional systems make it difficult for users to concretely express the vague image they have in their mind, which makes it difficult to find appropriate products and information. Furthermore, customization is not easy when the desired product is not available on the market or when users are dissatisfied with the search results. Furthermore, users often struggle to solidify a specific image. A system that can solve these problems is needed.
[0858] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0859] In this invention, the server includes: means for a user to input an image in their mind in language; means for the server to receive the input and generate an image based on the language data using a generative AI model; means for the server to perform an image search using the generated image to search for related products and information; means for a terminal to display the search results to the user; means for the terminal to call an AI concierge function when the user has difficulty solidifying a specific image; means for the server to pass the consultation content to the AI concierge function and propose an optimal image based on the user's wishes; and means for the terminal to display the proposal to the user. This allows users to concretely express vague images and, based on that, to search for, display, and customize information and products, and even to utilize requests for commercialization.
[0860] A "user" is someone who uses the system to input a mental image and search for products or information based on that image.
[0861] A "terminal" is a device that a user uses to input mental images as text and that displays search results and suggestions.
[0862] A "server" is a device or system that receives data sent by a user, generates images using a generative AI model, performs an image search, and returns the results to the terminal.
[0863] A "generative AI model" is an artificial intelligence model that analyzes text data entered by a user and generates visual images based on it.
[0864] A "prompt sentence" is a textual description that the user enters to concretely express the image in their mind.
[0865] An "image search engine" is a search engine that searches for related products and information based on images generated by the server.
[0866] "AI Concierge" is an artificial intelligence system that generates optimal suggestions when users are struggling with a specific choice.
[0867] A "customization request" is a user's request for further specific changes or improvements to the search results.
[0868] A "commercialization request" is a request for a product that a user desires to be produced or provided when the product does not exist on the market.
[0869] This invention is a system that allows users to concretely express the image in their mind and find appropriate information and products based on that image. This system consists of a user's device, a server, a generative AI model, and an image search engine.
[0870] User Input Phase
[0871] The user uses the device to input the image in their mind as text. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server. The device used can be a smartphone, tablet, PC, or other device, and the input field is provided in the form of a web application or mobile application.
[0872] Image generation phase
[0873] The server receives text data sent by the user. The server passes the received text data to a generative AI model (e.g., DALL-E or a similar generative model) to generate an image based on the text data. The generative AI model is a machine learning model that analyzes the text data and generates a visual image based on it. The generated image is stored on the server.
[0874] Image search phase
[0875] The generated image is sent from the server to an image search engine (for example, Google Image Search API) to search for related products and information. The server receives the search results returned from the image search engine and formats them in a user-friendly format.
[0876] Result display phase
[0877] The server then sends the formatted search results to the user's device, which then displays the received search results for the user to review. This allows users to easily find the products and information they want.
[0878] Customization Phase
[0879] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[0880] Commercialization request phase
[0881] If the product desired by the user does not exist on the market, the server will link the request to the relevant company. The server will then send the user's request and the generated image to the company, which will use it as a starting point to consider commercializing the product.
[0882] AI Concierge Phase
[0883] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to an AI concierge (e.g., GPT-4), which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0884] Specific examples
[0885] For example, if a user imagines a "retro-themed coffee shop interior," they can enter this text into their device and send it to the server. The server then uses a generative AI model to generate related images and collects information on suitable interiors through an image search engine. The results are then displayed on the device, allowing the user to view specific interior images and products.
[0886] Prompt Sentence Examples
[0887] "Think of lightweight sandals for the beach in the summer and search for related products."
[0888] "Imagine the interior of a retro coffee shop and propose a design that relates to that."
[0889] "Can you suggest an idea for a dress to wear to a colleague's wedding?"
[0890] By inputting a specific prompt sentence in this way, the system can provide appropriate information based on the user's wishes.
[0891] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0892] Step 1:
[0893] The user inputs text using the device. The user inputs the image in their mind as concrete text into the input field of the device. For example, the user might input "lightweight sandals to wear on the beach in the summer." The input data is sent from the device to the server in JSON format or HTTP request format.
[0894] Step 2:
[0895] The device sends the input data to the server as an HTTP POST request, with the text data included as the payload, which sends the image of the user's input to the server.
[0896] Step 3:
[0897] The server receives the text data. The server extracts the text data from the received HTTP request and prepares it for analysis. It receives text as input data and prepares the data to be passed to a generative AI model as output.
[0898] Step 4:
[0899] The server sends data to the generative AI model. The server creates and sends an API request to pass text data to the generative AI model (e.g., DALL-E). It sends text data as input and receives the generated image as output.
[0900] Step 5:
[0901] The AI model generates an image. The generative AI model analyzes the received text data and generates a visual image based on it. It receives text data as input and generates an image as output. The generated image is sent back to the server.
[0902] Step 6:
[0903] The server stores the generated image. The server receives the generated image and stores it in a database or file system. The image is stored in binary format for further processing.
[0904] Step 7:
[0905] The server sends the generated image to an image search engine. The server retrieves the stored image and sends it to an image search engine (e.g., Google Image Search API) in the form of a request, which sends the image as input and searches for related products and information.
[0906] Step 8:
[0907] The image search engine returns the search results. The image search engine searches for related products and information based on the received image and returns the results to the server. It takes an image as input and generates related search results as output.
[0908] Step 9:
[0909] The server formats the search results. The server converts the raw data received from the image search engine into an easy-to-understand format. It formats the product name, price, image URL, etc., and processes them into a format that can be presented to the user.
[0910] Step 10:
[0911] The server returns the formatted results to the terminal. The server then sends the formatted search results to the terminal as an HTTP response. The formatted data is sent and the results are provided to the user.
[0912] Step 11:
[0913] The device displays the search results it has received. The device analyzes the formatted search results it has received and displays them in the user interface in an appropriate layout. The user can then check products and information based on the results.
[0914] (Application example 1)
[0915] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0916] In the past, it was difficult for users to search for products or information by associating the mental image they had in their minds, which they could not specifically describe in words. Furthermore, there was a lack of support for making vague images concrete, making it difficult for users to properly find the products they wanted. To solve this problem, a system was needed that could visually materialize the image users have in their minds and effectively search for products and information based on that image.
[0917] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0918] In this invention, the server includes means for a user to input an image in their mind in language, means for generating an image based on the language data using a generative AI model, means for performing an image search using the generated image to search for related products and information, means for a user to input additional customization requests from a terminal, means for generating a customized image again using the generative AI model, and means for performing an image search using the customized image again. This makes it easier for a user to express a specific image, and enables products and information to be searched for and displayed based on that image.
[0919] "A means for users to input images in their minds in language" refers to an interface that allows users to input vague images or concepts in their minds as text.
[0920] "Generative AI model" refers to a machine learning model for automatically generating corresponding visual images based on specific text input.
[0921] "Image" refers to a visual representation generated from text by a generative AI model.
[0922] "Means for performing image search" refers to a function for searching for related products and information on the Internet based on the generated image.
[0923] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.
[0924] "Means for inputting customization requests" refers to an interface that allows the user to input more specific requests and conditions.
[0925] Generating a "customized image" refers to creating a new image using a generative AI model based on the customization request.
[0926] "Means for linking with companies" refers to a communication function for automatically notifying related companies of user requests and generated images.
[0927] "Commercialization methods" refers to the process by which a company analyzes or plans the manufacture and sale of a new product based on user requests.
[0928] "AI concierge function" refers to an artificial intelligence-based support function that provides optimal images and suggestions based on the user's vague requests and questions.
[0929] This invention is a system that allows users to input text images that they cannot express in concrete form, and find appropriate information and products based on those images. The system consists of a user terminal, a server, a generative AI model, and an image search engine.
[0930] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server.
[0931] The server receives text data sent by the user. The server passes the text data to a generative AI model, which generates an image. Specifically, the generative AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server. Specifically, open-source generative AI models such as DALL-E and Imagen can be applied.
[0932] The generated image is sent from the server to an image search engine, which searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format. For this purpose, AWS Rekognition or Google Cloud Vision API can be used, for example.
[0933] The server sends the formatted search results to the user's terminal, which displays the received search results so that the user can check them.
[0934] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request on their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search, and the results are provided to the user.
[0935] Furthermore, if the product desired by the user does not exist, the server can link the request to relevant companies. The server then sends the user's request and the generated image to the company, providing a starting point for the company to consider commercializing the product.
[0936] Furthermore, if a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[0937] As a concrete example, consider the case where a user launches an application and enters "warm boots that won't slip on snowy winter roads." The server receives the text, and the generative AI model generates an image of "warm boots that won't slip on snowy winter roads." That image is then run through an image search engine, and related products are displayed. If the user further requests customization by adding "red boots," another customized image is generated and the results are displayed.
[0938] An example prompt is, "A user is searching for warm boots that will keep them from slipping on snowy winter roads. Please process this text input in the following format and generate an image using a generative AI model."
[0939] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0940] Step 1:
[0941] The user uses a terminal to input text representing the image in their mind. This input data is a specific description such as "Warm boots that won't slip on snowy roads in winter." This text input is sent from the user terminal to the server.
[0942] Step 2:
[0943] The server receives text data sent by the user (input). It then analyzes the text using a generative AI model and generates an image based on the input language data (data processing and data calculation). The generated image is saved on the server (output). Specifically, generative AI models such as DALL-E and Imagen are used.
[0944] Step 3:
[0945] The server sends the generated image to an image search engine (input). The image search engine searches for related products and information based on the sent image (data processing and data calculation). The search results are returned to the server (output), which formats them in a form that is easy for users to understand. Specifically, AWS Rekognition and Google Cloud Vision API are used.
[0946] Step 4:
[0947] The server sends the formatted search results to the user's terminal (input). The user's terminal receives the search results and displays them to the user (output). The user can check the displayed results.
[0948] Step 5:
[0949] If the user is not satisfied with the search results, they can input additional customization requests (input). For example, they can input a specific request such as "more brightly colored sandals." This customization request is sent from the terminal to the server.
[0950] Step 6:
[0951] The server receives the customization request (input). A new customized image is generated using the generative AI model again (data processing and data calculation). The generated image is saved on the server (output).
[0952] Step 7:
[0953] The server performs another image search using the customized image (input). The image search engine searches for related products and information based on the new image (data processing and data calculation). The search results are returned to the server (output) and reformatted again.
[0954] Step 8:
[0955] The server sends the formatted customized search results to the user's device (input). The user's device receives the search results again and displays them (output). This allows the user to check more specific products and information.
[0956] Step 9:
[0957] If the product desired by the user is not available, the user's request and the generated image are notified to related companies (input). The server then shares this information with interested companies, who then consider commercializing the product (output).
[0958] Step 10:
[0959] If the user has difficulty solidifying a specific image, they can call the AI concierge function from their device (input). The server receives the consultation details, passes them to the AI concierge, and proposes the optimal image (data processing and data calculation). The proposed image is then displayed on the user's device (output).
[0960] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0961] This invention is a system that helps users concretely express the images in their minds and provides appropriate information and products taking into account the user's emotions. This system is composed of a user terminal, a server, an AI model, an emotion engine, and an image search engine.
[0962] User Input Phase
[0963] The user uses the device to input text representing the image in their mind, for example, "lightweight sandals to wear on the beach in the summer." As the user inputs text, the device captures the input text data in real time and sends it to the server.
[0964] Emotion Recognition Phase
[0965] The server receives the text data sent from the device and passes it to the emotion engine. The emotion engine analyzes the text data and recognizes the user's emotion. For example, if the user imagines a "fun beach vacation," the emotion engine recognizes "fun."
[0966] Image generation phase
[0967] The server requests the AI model to generate an image based on the emotion data from the emotion engine. Based on the input text and emotion data, the AI model generates an image of, for example, "summer beach sandals" in bright colors that reflect "fun." The generated image is stored on the server.
[0968] Image search phase
[0969] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[0970] Result display phase
[0971] The server returns an API response that sends the formatted search results to the user's device. The device displays the received search results to the user, who can then review them to find the products or information they are looking for.
[0972] Customization Phase (Optional)
[0973] If the user is not satisfied with the search results, they can input additional customization requests. For example, if the user wants "sandals in a more vibrant color," they input that request from their device and send it to the server. The server receives the customization request and generates a new image using the emotion engine and AI model. For example, it generates a new image of "sandals" that reflects "fun" in a more vibrant color. It then runs that image through an image search again and provides the results to the user.
[0974] Commercialization request phase (optional)
[0975] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[0976] AI Concierge Phase (Optional)
[0977] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge and requests suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will suggest the most suitable outfit image based on that. These suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on them.
[0978] As described above, the present invention allows users to concretely express the images and emotions in their minds, and based on those images, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization. This allows users to access products and services that match their emotions and desires.
[0979] The processing flow will be explained below.
[0980] Step 1:
[0981] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[0982] Step 2:
[0983] The device captures the user's input text in real time and makes an API request to send that data to the server.
[0984] Step 3:
[0985] The server receives the user's input text sent from the device via the API.
[0986] Step 4:
[0987] The server passes the received text data to the emotion engine to recognize the user's emotion.
[0988] Step 5:
[0989] The emotion engine analyzes the text data and identifies the user's emotion, for example, recognizing the emotion that indicates "fun."
[0990] Step 6:
[0991] The server passes the recognized emotion data to the AI model and makes an image generation request.
[0992] Step 7:
[0993] Based on the input text and emotional data, the AI model generates an image of "summer beach sandals" that reflects "fun."
[0994] Step 8:
[0995] The server receives the images generated by the AI model and stores them in a database.
[0996] Step 9:
[0997] The image stored by the server is sent to an image search engine to search for related products and information.
[0998] Step 10:
[0999] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[1000] Step 11:
[1001] The server returns the formatted search results as an API response to be sent to the terminal.
[1002] Step 12:
[1003] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[1004] Step 13:
[1005] If the user is not satisfied with the search results, he or she can input additional customization requests. For example, if the user desires "sandals in brighter colors," the request is input from the terminal and transmitted to the server.
[1006] Step 14:
[1007] The device obtains the user's customization requests and makes an API request to send them back to the server.
[1008] Step 15:
[1009] The server receives the customization request and passes it back to the emotion engine to reconfirm the user's emotion.
[1010] Step 16:
[1011] The emotion engine analyzes the text data included in the customization request and recognizes the user's emotion.
[1012] Step 17:
[1013] The server again passes the recognized emotion data to the AI model and makes a customized image generation request.
[1014] Step 18:
[1015] The AI model generates new images based on customized criteria and emotional data.
[1016] Step 19:
[1017] The server receives the new image and sends it back to the image search engine.
[1018] Step 20:
[1019] The server receives the search results from the image search engine again, formats them, and sends them back to the user's terminal.
[1020] Step 21:
[1021] The terminal receives the customized search results from the server and displays them to the user.
[1022] Step 22:
[1023] If the product desired by the user is not available on the market, the terminal notifies the server of this fact.
[1024] Step 23:
[1025] The server sends the user's request and the generated image to a company, which considers commercializing the image.
[1026] Step 24:
[1027] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[1028] Step 25:
[1029] The server passes the consultation content and emotional data to the AI concierge and requests a proposal.
[1030] Step 26:
[1031] The AI concierge generates optimal suggestions based on input text and sentiment data.
[1032] Step 27:
[1033] The server receives the AI concierge's suggestions and sends them to the user's device.
[1034] Step 28:
[1035] The terminal displays the suggestions to the user, who can then make an appropriate selection.
[1036] Example 2
[1037] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1038] In the past, it was difficult for users to find a specific product that matches their mental image, and there were no systems that provided products or information that took the user's emotions into consideration. This made it difficult to provide products that matched the user's needs and emotions, and improving user satisfaction was a challenge.
[1039] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1040] In this invention, the server includes means for a user to input the image in their mind in language, means for receiving the input and analyzing the language data using an emotion recognition engine to recognize the user's emotion, means for the server to generate an image from the language data using a generative AI model based on the recognized emotion data, means for the server to use the generated image to search for related products and information using an image search engine, and means for a terminal to display the search results to the user. This makes it possible to quickly and accurately provide products and information based on the user's needs and emotions.
[1041] "User" refers to an individual or group that uses the system to materialize the images and desires in their mind.
[1042] "Terminal" refers to an input device or display device used by a user, including a personal computer, smartphone, tablet, etc.
[1043] "Server" refers to the computer system responsible for data processing, data storage, and execution of various engines and models for the entire system.
[1044] "Language data" refers to text information entered by the user, and is made up of sentences and words that express the user's mental images and wishes.
[1045] An "emotion recognition engine" refers to an algorithm or program that analyzes the language data entered by a user and identifies the user's emotions from it.
[1046] A "generative AI model" refers to an artificial intelligence model that generates output such as images or text based on input data, and performs creative generation in response to specific prompts.
[1047] A "prompt sentence" is text data input into a generative AI model, and refers to a sentence that contains instructions for generating the output image or information.
[1048] "Image" refers to a visual representation generated by a generative AI model based on the user's language and emotional data.
[1049] An "image search engine" refers to a system or program that searches the Internet for related products and information based on an input image and returns the results.
[1050] "Search results" refers to the list of related products and information returned by an image search engine, including relevant data that may be useful to the user.
[1051] "Customization request" refers to content that the user wishes to add or change based on the search results, and includes new input for regenerating or searching.
[1052] "Company" refers to a corporation or organization that considers commercialization in response to user requests.
[1053] This invention is a system that allows users to concretely express the images in their minds and provides appropriate information and products by taking into account the user's emotions. This system is composed of a user terminal, a server, an emotion recognition engine, a generative AI model, and an image search engine.
[1054] User Input Phase
[1055] The user uses the device to input the image in their mind in words. For example, they might input "lightweight sandals to wear on the beach in the summer." The device captures the text data entered by the user in real time and sends it to a server. The device can be a PC, smartphone, tablet, or other device.
[1056] Emotion Recognition Phase
[1057] The server receives the text data sent from the device and passes it to an emotion recognition engine. The emotion recognition engine (such as "EmotionAPI") analyzes the text data and recognizes the user's emotions. For example, if the user imagines a "fun beach vacation," the emotion recognition engine will identify the emotion "enjoyment."
[1058] Image generation phase
[1059] The server requests the generative AI model to generate an image based on the emotion data obtained from the emotion recognition engine. The generative AI model (e.g., "DALL-E") generates an image based on the input text and emotion data.
[1060] Examples of prompts include:
[1061] "Lightweight sandals perfect for summer beach days. Brightly colored designs for a fun beach vacation."
[1062] Based on this, the generative AI model generates a brightly colored image of "summer flip-flops," which is then stored on a server.
[1063] Image search phase
[1064] The generated image is sent from the server to an image search engine. The image search engine (for example, Google Image Search API) searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format.
[1065] Result display phase
[1066] The server sends the formatted search results to the user's device, which then displays them to the user, who can then review them to find the products or information they want.
[1067] Customization Phase (Optional)
[1068] If a user is not satisfied with the search results, they can input additional customization requests. For example, if they want "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again generates a new image using the emotion recognition engine and generative AI model. The generated image is then sent to the image search engine again, and the results are provided to the user.
[1069] Commercialization request phase (optional)
[1070] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1071] AI Concierge Phase (Optional)
[1072] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if they are unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge, which then makes optimal suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will use that information to suggest the most appropriate outfit. These suggestions are displayed on the user's device, allowing them to make an appropriate choice.
[1073] As a result, the present invention can provide products and information that are in line with the user's feelings and desires more accurately, thereby improving user satisfaction.
[1074] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1075] Step 1: User Input Phase
[1076] The user inputs the image in their mind using language. For example, they might input "lightweight sandals to wear on the beach in the summer."
[1077] The input language data is acquired in real time by the terminal, and the terminal transmits the acquired text data to the server.
[1078] Input: User text input (e.g., "Lightweight sandals for summer beach wear")
[1079] Output: Send text data from the terminal to the server
[1080] Step 2: Emotion Recognition Phase
[1081] The server passes the text data received from the device to the emotion recognition engine, which analyzes the text data and recognizes the user's emotions.
[1082] The text "Lightweight sandals to wear on the beach in summer" is passed to an emotion recognition engine, which analyzes it and extracts the emotion data "fun."
[1083] Input: Text data received by the server
[1084] Output: Emotion data extracted by the emotion recognition engine (e.g., "enjoyment")
[1085] Step 3: Image generation phase
[1086] The server sends a prompt to the generative AI model based on the emotion data obtained from the emotion recognition engine, requesting image generation. The generative AI model generates an image based on the input text and emotion data.
[1087] Prompt: "Lightweight sandals for summer beach wear. Brightly colored designs for a fun beach vacation."
[1088] Based on this prompt, the generative AI model generates a brightly colored image of "summer flip-flops" and stores the image on a server.
[1089] Input: Prompt text and emotion data sent by the server
[1090] Output: Generated image
[1091] Step 4: Image Search Phase
[1092] The generated image is sent from the server to an image search engine, which searches for related products and information and returns the search results to the server.
[1093] The server formats the received search results in a user-friendly format.
[1094] Input: Image sent by the server
[1095] Output: Related product information returned by an image search engine
[1096] Step 5: Result display phase
[1097] The server sends the formatted search results to the user's terminal, which then displays the received search results to the user.
[1098] The user checks and selects the products and information they want from the displayed search results.
[1099] Input: The formatted search results sent by the server
[1100] Output: Search results displayed on your terminal
[1101] Step 6: Customization Phase (Optional)
[1102] If the user is not satisfied with the search results, they can input additional customization requests, such as "more vibrantly colored sandals," by entering the request on their device and sending it to the server.
[1103] The server receives the customization request, generates a new image using the emotion recognition engine and generative AI model, and then sends the generated image to the image search engine again, providing the search results to the user.
[1104] Input: Customization requests entered by the user
[1105] Output: Regenerated image and new search results
[1106] Step 7: Commercialization Request Phase (Optional)
[1107] If the product desired by the user does not exist on the market, the server will link this request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1108] Input: Product development requests entered by the user
[1109] Output: Requests and images sent to the company
[1110] Step 8: AI Concierge Phase (Optional)
[1111] If the user has difficulty solidifying a specific image, they can call up the AI concierge function on their device and input their inquiry. For example, if they are unsure what to wear to a colleague's wedding, they can input their inquiry into their device.
[1112] The server acquires the consultation details and emotional data and passes them to the AI concierge. The AI concierge then makes optimal suggestions based on this information. For example, if the user is feeling "joy" or "elegance," it will suggest the perfect outfit. The suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on this information.
[1113] Input: Consultation content entered by the user
[1114] Output: Suggestions from the AI concierge
[1115] (Application example 2)
[1116] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1117] Conventional systems make it difficult for users to search for products based on specific product images or emotions, making it difficult to find products that fully meet the user's needs. Furthermore, with conventional text searches that are not based on emotions, there can be a gap between the products the user is looking for and the search results. The present invention aims to solve these problems by providing a more accurate product search system based on user emotions and text.
[1118] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input the image in their mind in language, means for recognizing the user's emotion from the language data using an emotion engine, means for generating an image based on the language data and emotion data using an AI model, means for the server to perform an image search using the generated image to search for related products and information, and means for the terminal to display the search results to the user. This makes it possible to quickly provide products and information based on the user's emotions and specific needs.
[1119] A "user" is an individual or entity that uses the system to accomplish a particular task.
[1120] A "mental image" is a visual representation of a particular scene or object that a user has in their mind.
[1121] "Language data" is text data that a user inputs to express an image in their mind.
[1122] A "server" is a computer system on a network that receives and processes data sent from user terminals.
[1123] An "emotion engine" is a software system that analyzes a user's language data and recognizes the emotions contained therein.
[1124] "Emotion data" is emotion information extracted from the user's language data analyzed by the emotion engine.
[1125] An "AI model" is a mathematical model that uses artificial intelligence technology to perform specific tasks.
[1126] An "image" is a visual image generated by an AI model based on a user's language and emotional data.
[1127] "Image search" is the process of searching the Internet for related products and information based on a generated image.
[1128] "Related products and information" refers to items and data found through image search based on the user's needs and emotions.
[1129] "Search results" are lists of related products and information obtained through image search.
[1130] A "terminal" is a device that can be directly operated by a user, such as a smartphone or computer.
[1131] A "customization request" is a request for improvement that a user inputs when the user is dissatisfied with the search results.
[1132] A "customized image" is an image generated again by an AI model based on customization requests.
[1133] "Interested companies" are organizations that have the potential to respond to user requests and develop or improve products.
[1134] "Marketing information" is information that includes user emotion data that companies can use as a reference when considering commercialization.
[1135] To implement this invention, the user must first input the image in their mind as concrete text. When the user inputs the text using a device such as a smartphone or computer, the data is sent to a server.
[1136] The server uses an emotion engine, such as Hugging Face's "sentiment-analysis" transformer model, to analyze the user's emotions from the received text data. The emotion engine analyzes the words and expressions in the text data and can recognize the emotions conveyed in the user's input in real time.
[1137] The server then combines the emotion data obtained from the emotion engine with the text data and generates an image using a generative AI model such as DALL-E. This is the process of generating a visual image based on the text entered by the user, further reflecting the emotion contained in the text. The generated image is then stored on the server.
[1138] The generated images are used to search for related products and information using image search engines such as the Google Image Search API. The server receives the search results returned by the image search engine and formats them in a user-friendly format.
[1139] The server sends the formatted search results to the user's device, which then displays them to the user. The user can then review the products and information displayed in the search results and make an appropriate selection.
[1140] If the user is not satisfied with the search results, they can input additional customization requests from their device. The server receives these customization requests and generates customized images using the emotion engine and generative AI model again. Further image searches are performed using these customized images, and the results are displayed to the user.
[1141] If the product desired by the user is not available on the market, the server will share the request with relevant companies. At this time, the generated image and emotion data will also be provided to the companies. Based on this, the companies can consider commercializing the product and analyze marketing information.
[1142] As a concrete example, consider a user searching for "swimsuits for a fun weekend at the beach." When the user enters this text, the emotion engine recognizes the word "fun," and the generative AI model generates a bright and cheerful image of the swimsuit. The server then searches for related products based on this image and suggests them to the user.
[1143] Example prompt sentence:
[1144] Your goal is to implement an application that recognizes emotions from the input text, such as "I'm looking for a swimsuit for a fun weekend at the beach," generates relevant images based on the emotion, and searches for and suggests related products.
[1145] In this way, the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[1146] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1147] Step 1:
[1148] The user inputs the image in their mind using language. For example, if they are looking for a swimsuit for a fun weekend at the beach, they can input this text using a device such as a smartphone or computer. The input data (text) is acquired in real time and sent to the server.
[1149] Step 2:
[1150] The server receives the text data sent from the device. Then, it uses an emotion engine (for example, Hugging Face's "sentiment-analysis" Transformer model) to analyze the user's emotions from the received text data. In this case, the input is the user's text data, and the output is emotional data such as "enjoyment."
[1151] Step 3:
[1152] The server requests a generative AI model (such as DALL-E) to generate an image based on the emotion data and text data from the emotion engine. The input is text data and emotion data, and the output is an image that reflects the emotion. Specifically, the server generates an image of a bright and cheerful beach swimsuit that reflects "fun."
[1153] Step 4:
[1154] Using the generated image, the server uses an image search engine (such as the Google Image Search API) to search for related products and information. In this case, the input is the image, and the output is a list of related products and information.
[1155] Step 5:
[1156] The server receives the search results returned by the image search engine and formats them in a user-friendly format. This formatting process ensures that the search results are displayed in a format that is appealing to the user. The input to this process is the unformatted search result data returned by the search engine, and the output is the formatted search result data.
[1157] Step 6:
[1158] The server sends the formatted search results to the user's device, which then displays the received search results to the user. The input here is the formatted search result data, and the output is the visual search results displayed on the user's device screen.
[1159] Step 7:
[1160] If the user is not satisfied with a particular product or information, they can input additional customization requests from their terminal. These requests are sent back to the server and used for the next process. The input is the text data of the user's customization requests, and the output is the customization requests sent to the server.
[1161] Step 8:
[1162] The server receives the customization request and generates a customized image using the emotion engine and generative AI model. Here, the input is the text data and emotion data of the customization request, and the output is a customized image. For example, if a user requests "sandals with a more vibrant color," the server generates an image of more vibrant sandals.
[1163] Step 9:
[1164] The server then performs another image search using the customized image, where the input is the customized image and the output is a new list of related products and information, formats the search results, and presents them to the user again.
[1165] Step 10:
[1166] If the product desired by the user does not exist on the market, the server will share this request with relevant companies. The companies will then consider commercializing the product. The input here is the user's request, the generated image, and emotion data, and the output is marketing information provided to the company and the results of consideration for commercialization.
[1167] In this way, the system of the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[1168] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1169] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1170] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1171] [Fourth embodiment]
[1172] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1173] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1174] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1175] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1176] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1177] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1178] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1179] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1180] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1181] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1182] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1183] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1184] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1185] This invention is a system that allows users to input images in their minds that they cannot express concretely in words, and find appropriate information and products based on those images. The system consists of a user terminal, a server, an AI model, and an image search engine.
[1186] User Input Phase
[1187] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the device's input field. This input data is sent from the device to the server.
[1188] Image generation phase
[1189] The server receives the text data sent by the user. The server passes the received text data to the AI model, which generates an image. Specifically, the AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server.
[1190] Image search phase
[1191] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[1192] Result display phase
[1193] The server sends the formatted search results to the user's device, where the user can view and check the results. This allows the user to easily find the products and information they want.
[1194] Customization Phase (Optional)
[1195] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[1196] Commercialization request phase (optional)
[1197] If the product desired by the user does not exist on the market, the server will link the request to relevant companies. The server will then send the user's request and the generated image to the companies, which will then use it as a starting point for considering commercialization.
[1198] AI Concierge Phase (Optional)
[1199] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[1200] As described above, the present invention allows users to concretely express the vague image in their mind, and based on that, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization.
[1201] The processing flow will be explained below.
[1202] Step 1:
[1203] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[1204] Step 2:
[1205] The device captures the user's input text in real time and makes an API request to send that data to the server.
[1206] Step 3:
[1207] The server receives the user's input text sent from the device via the API.
[1208] Step 4:
[1209] The server passes the received text data to the AI model and requests it to generate an image. The AI model analyzes the input text and generates a corresponding image.
[1210] Step 5:
[1211] The server receives the images generated by the AI model and stores them in a database.
[1212] Step 6:
[1213] The image stored by the server is sent to an image search engine to search for related products and information.
[1214] Step 7:
[1215] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[1216] Step 8:
[1217] The server returns an API response that sends the formatted search results to the terminal.
[1218] Step 9:
[1219] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[1220] Step 10:
[1221] If the user is not satisfied with the search results, they can enter their request for further customization into the device's input field, for example, "more brightly colored sandals."
[1222] Step 11:
[1223] The device obtains the user's customization requests and makes an API request to send them back to the server.
[1224] Step 12:
[1225] The server receives the customization request and requests the AI model to generate an image again. The AI model generates a new image based on the customized conditions.
[1226] Step 13:
[1227] The server receives the new image and causes the image search engine to search again.
[1228] Step 14:
[1229] The server receives the new search results, formats them, and sends them back to the user's device, which then displays the customized search results to the user.
[1230] Step 15:
[1231] If the product desired by the user does not exist on the market, the server will connect the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1232] Step 16:
[1233] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[1234] Step 17:
[1235] The server passes the consultation details to the AI concierge and requests a proposal. The AI concierge analyzes the consultation details and proposes the optimal image.
[1236] Step 18:
[1237] The server receives the AI concierge's suggestions and sends them to the user's device, which displays the suggested images to the user, allowing the user to make an appropriate choice.
[1238] Example 1
[1239] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1240] Conventional systems make it difficult for users to concretely express the vague image they have in their mind, which makes it difficult to find appropriate products and information. Furthermore, customization is not easy when the desired product is not available on the market or when users are dissatisfied with the search results. Furthermore, users often struggle to solidify a specific image. A system that can solve these problems is needed.
[1241] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1242] In this invention, the server includes: means for a user to input an image in their mind in language; means for the server to receive the input and generate an image based on the language data using a generative AI model; means for the server to perform an image search using the generated image to search for related products and information; means for a terminal to display the search results to the user; means for the terminal to call an AI concierge function when the user has difficulty solidifying a specific image; means for the server to pass the consultation content to the AI concierge function and propose an optimal image based on the user's wishes; and means for the terminal to display the proposal to the user. This allows users to concretely express vague images and, based on that, to search for, display, and customize information and products, and even to utilize requests for commercialization.
[1243] A "user" is someone who uses the system to input a mental image and search for products or information based on that image.
[1244] A "terminal" is a device that a user uses to input mental images as text and that displays search results and suggestions.
[1245] A "server" is a device or system that receives data sent by a user, generates images using a generative AI model, performs an image search, and returns the results to the terminal.
[1246] A "generative AI model" is an artificial intelligence model that analyzes text data entered by a user and generates visual images based on it.
[1247] A "prompt sentence" is a textual description that the user enters to concretely express the image in their mind.
[1248] An "image search engine" is a search engine that searches for related products and information based on images generated by the server.
[1249] "AI Concierge" is an artificial intelligence system that generates optimal suggestions when users are struggling with a specific choice.
[1250] A "customization request" is a user's request for further specific changes or improvements to the search results.
[1251] A "commercialization request" is a request for a product that a user desires to be produced or provided when the product does not exist on the market.
[1252] This invention is a system that allows users to concretely express the image in their mind and find appropriate information and products based on that image. This system consists of a user's device, a server, a generative AI model, and an image search engine.
[1253] User Input Phase
[1254] The user uses the device to input the image in their mind as text. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server. The device used can be a smartphone, tablet, PC, or other device, and the input field is provided in the form of a web application or mobile application.
[1255] Image generation phase
[1256] The server receives text data sent by the user. The server passes the received text data to a generative AI model (e.g., DALL-E or a similar generative model) to generate an image based on the text data. The generative AI model is a machine learning model that analyzes the text data and generates a visual image based on it. The generated image is stored on the server.
[1257] Image search phase
[1258] The generated image is sent from the server to an image search engine (for example, Google Image Search API) to search for related products and information. The server receives the search results returned from the image search engine and formats them in a user-friendly format.
[1259] Result display phase
[1260] The server then sends the formatted search results to the user's device, which then displays the received search results for the user to review. This allows users to easily find the products and information they want.
[1261] Customization Phase
[1262] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search and the results are provided to the user.
[1263] Commercialization request phase
[1264] If the product desired by the user does not exist on the market, the server will link the request to the relevant company. The server will then send the user's request and the generated image to the company, which will use it as a starting point to consider commercializing the product.
[1265] AI Concierge Phase
[1266] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to an AI concierge (e.g., GPT-4), which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[1267] Specific examples
[1268] For example, if a user imagines a "retro-themed coffee shop interior," they can enter this text into their device and send it to the server. The server then uses a generative AI model to generate related images and collects information on suitable interiors through an image search engine. The results are then displayed on the device, allowing the user to view specific interior images and products.
[1269] Prompt Sentence Examples
[1270] "Think of lightweight sandals for the beach in the summer and search for related products."
[1271] "Imagine the interior of a retro coffee shop and propose a design that relates to that."
[1272] "Can you suggest an idea for a dress to wear to a colleague's wedding?"
[1273] By inputting a specific prompt sentence in this way, the system can provide appropriate information based on the user's wishes.
[1274] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1275] Step 1:
[1276] The user inputs text using the device. The user inputs the image in their mind as concrete text into the input field of the device. For example, the user might input "lightweight sandals to wear on the beach in the summer." The input data is sent from the device to the server in JSON format or HTTP request format.
[1277] Step 2:
[1278] The device sends the input data to the server as an HTTP POST request, with the text data included as the payload, which sends the image of the user's input to the server.
[1279] Step 3:
[1280] The server receives the text data. The server extracts the text data from the received HTTP request and prepares it for analysis. It receives text as input data and prepares the data to be passed to a generative AI model as output.
[1281] Step 4:
[1282] The server sends data to the generative AI model. The server creates and sends an API request to pass text data to the generative AI model (e.g., DALL-E). It sends text data as input and receives the generated image as output.
[1283] Step 5:
[1284] The AI model generates an image. The generative AI model analyzes the received text data and generates a visual image based on it. It receives text data as input and generates an image as output. The generated image is sent back to the server.
[1285] Step 6:
[1286] The server stores the generated image. The server receives the generated image and stores it in a database or file system. The image is stored in binary format for further processing.
[1287] Step 7:
[1288] The server sends the generated image to an image search engine. The server retrieves the stored image and sends it to an image search engine (e.g., Google Image Search API) in the form of a request, which sends the image as input and searches for related products and information.
[1289] Step 8:
[1290] The image search engine returns the search results. The image search engine searches for related products and information based on the received image and returns the results to the server. It takes an image as input and generates related search results as output.
[1291] Step 9:
[1292] The server formats the search results. The server converts the raw data received from the image search engine into an easy-to-understand format. It formats the product name, price, image URL, etc., and processes them into a format that can be presented to the user.
[1293] Step 10:
[1294] The server returns the formatted results to the terminal. The server then sends the formatted search results to the terminal as an HTTP response. The formatted data is sent and the results are provided to the user.
[1295] Step 11:
[1296] The device displays the search results it has received. The device analyzes the formatted search results it has received and displays them in the user interface in an appropriate layout. The user can then check products and information based on the results.
[1297] (Application example 1)
[1298] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1299] In the past, it was difficult for users to search for products or information by associating the mental image they had in their minds, which they could not specifically describe in words. Furthermore, there was a lack of support for making vague images concrete, making it difficult for users to properly find the products they wanted. To solve this problem, a system was needed that could visually materialize the image users have in their minds and effectively search for products and information based on that image.
[1300] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1301] In this invention, the server includes means for a user to input an image in their mind in language, means for generating an image based on the language data using a generative AI model, means for performing an image search using the generated image to search for related products and information, means for a user to input additional customization requests from a terminal, means for generating a customized image again using the generative AI model, and means for performing an image search using the customized image again. This makes it easier for a user to express a specific image, and enables products and information to be searched for and displayed based on that image.
[1302] "A means for users to input images in their minds in language" refers to an interface that allows users to input vague images or concepts in their minds as text.
[1303] "Generative AI model" refers to a machine learning model for automatically generating corresponding visual images based on specific text input.
[1304] "Image" refers to a visual representation generated from text by a generative AI model.
[1305] "Means for performing image search" refers to a function for searching for related products and information on the Internet based on the generated image.
[1306] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.
[1307] "Means for inputting customization requests" refers to an interface that allows the user to input more specific requests and conditions.
[1308] Generating a "customized image" refers to creating a new image using a generative AI model based on the customization request.
[1309] "Means for linking with companies" refers to a communication function for automatically notifying related companies of user requests and generated images.
[1310] "Commercialization methods" refers to the process by which a company analyzes or plans the manufacture and sale of a new product based on user requests.
[1311] "AI concierge function" refers to an artificial intelligence-based support function that provides optimal images and suggestions based on the user's vague requests and questions.
[1312] This invention is a system that allows users to input text images that they cannot express in concrete form, and find appropriate information and products based on those images. The system consists of a user terminal, a server, a generative AI model, and an image search engine.
[1313] The user uses the device to input text representing the image in their mind. For example, if the user is looking for "lightweight sandals to wear on the beach in the summer," they input that specific image as text into the input field on the device. This input data is sent from the device to the server.
[1314] The server receives text data sent by the user. The server passes the text data to a generative AI model, which generates an image. Specifically, the generative AI model analyzes the text "summer beach sandals" and generates a visual image based on it. This generated image is stored on the server. Specifically, open-source generative AI models such as DALL-E and Imagen can be applied.
[1315] The generated image is sent from the server to an image search engine, which searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format. For this purpose, AWS Rekognition or Google Cloud Vision API can be used, for example.
[1316] The server sends the formatted search results to the user's terminal, which displays the received search results so that the user can check them.
[1317] If a user is not satisfied with the search results, they can input additional customization requests. For example, if a user wants "sandals in a brighter color," they input that request on their device and send it to the server. The server receives the customization request and again uses the generative AI model to generate a customized image. The image is then subjected to another image search, and the results are provided to the user.
[1318] Furthermore, if the product desired by the user does not exist, the server can link the request to relevant companies. The server then sends the user's request and the generated image to the company, providing a starting point for the company to consider commercializing the product.
[1319] Furthermore, if a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their inquiry into their device. The server passes the inquiry to the AI concierge, which then suggests the most suitable outfit based on the user's preferences. This suggestion is displayed on the user's device, allowing the user to make an appropriate choice.
[1320] As a concrete example, consider the case where a user launches an application and enters "warm boots that won't slip on snowy winter roads." The server receives the text, and the generative AI model generates an image of "warm boots that won't slip on snowy winter roads." That image is then run through an image search engine, and related products are displayed. If the user further requests customization by adding "red boots," another customized image is generated and the results are displayed.
[1321] An example prompt is, "A user is searching for warm boots that will keep them from slipping on snowy winter roads. Please process this text input in the following format and generate an image using a generative AI model."
[1322] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1323] Step 1:
[1324] The user uses a terminal to input text representing the image in their mind. This input data is a specific description such as "Warm boots that won't slip on snowy roads in winter." This text input is sent from the user terminal to the server.
[1325] Step 2:
[1326] The server receives text data sent by the user (input). It then analyzes the text using a generative AI model and generates an image based on the input language data (data processing and data calculation). The generated image is saved on the server (output). Specifically, generative AI models such as DALL-E and Imagen are used.
[1327] Step 3:
[1328] The server sends the generated image to an image search engine (input). The image search engine searches for related products and information based on the sent image (data processing and data calculation). The search results are returned to the server (output), which formats them in a form that is easy for users to understand. Specifically, AWS Rekognition and Google Cloud Vision API are used.
[1329] Step 4:
[1330] The server sends the formatted search results to the user's terminal (input). The user's terminal receives the search results and displays them to the user (output). The user can check the displayed results.
[1331] Step 5:
[1332] If the user is not satisfied with the search results, they can input additional customization requests (input). For example, they can input a specific request such as "more brightly colored sandals." This customization request is sent from the terminal to the server.
[1333] Step 6:
[1334] The server receives the customization request (input). A new customized image is generated using the generative AI model again (data processing and data calculation). The generated image is saved on the server (output).
[1335] Step 7:
[1336] The server performs another image search using the customized image (input). The image search engine searches for related products and information based on the new image (data processing and data calculation). The search results are returned to the server (output) and reformatted again.
[1337] Step 8:
[1338] The server sends the formatted customized search results to the user's device (input). The user's device receives the search results again and displays them (output). This allows the user to check more specific products and information.
[1339] Step 9:
[1340] If the product desired by the user is not available, the user's request and the generated image are notified to related companies (input). The server then shares this information with interested companies, who then consider commercializing the product (output).
[1341] Step 10:
[1342] If the user has difficulty solidifying a specific image, they can call the AI concierge function from their device (input). The server receives the consultation details, passes them to the AI concierge, and proposes the optimal image (data processing and data calculation). The proposed image is then displayed on the user's device (output).
[1343] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1344] This invention is a system that helps users concretely express the images in their minds and provides appropriate information and products taking into account the user's emotions. This system is composed of a user terminal, a server, an AI model, an emotion engine, and an image search engine.
[1345] User Input Phase
[1346] The user uses the device to input text representing the image in their mind, for example, "lightweight sandals to wear on the beach in the summer." As the user inputs text, the device captures the input text data in real time and sends it to the server.
[1347] Emotion Recognition Phase
[1348] The server receives the text data sent from the device and passes it to the emotion engine. The emotion engine analyzes the text data and recognizes the user's emotion. For example, if the user imagines a "fun beach vacation," the emotion engine recognizes "fun."
[1349] Image generation phase
[1350] The server requests the AI model to generate an image based on the emotion data from the emotion engine. Based on the input text and emotion data, the AI model generates an image of, for example, "summer beach sandals" in bright colors that reflect "fun." The generated image is stored on the server.
[1351] Image search phase
[1352] The generated image is sent from the server to an image search engine, where related products and information are searched for. The server receives the search results from the image search engine and formats them in a user-friendly format.
[1353] Result display phase
[1354] The server returns an API response that sends the formatted search results to the user's device. The device displays the received search results to the user, who can then review them to find the products or information they are looking for.
[1355] Customization Phase (Optional)
[1356] If the user is not satisfied with the search results, they can input additional customization requests. For example, if the user wants "sandals in a more vibrant color," they input that request from their device and send it to the server. The server receives the customization request and generates a new image using the emotion engine and AI model. For example, it generates a new image of "sandals" that reflects "fun" in a more vibrant color. It then runs that image through an image search again and provides the results to the user.
[1357] Commercialization request phase (optional)
[1358] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1359] AI Concierge Phase (Optional)
[1360] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if a user is unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge and requests suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will suggest the most suitable outfit image based on that. These suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on them.
[1361] As described above, the present invention allows users to concretely express the images and emotions in their minds, and based on those images, realizes a series of processes that search for and display products and information, and even leads to requests for commercialization. This allows users to access products and services that match their emotions and desires.
[1362] The processing flow will be explained below.
[1363] Step 1:
[1364] The user enters a mental image into the device's input field as text, for example, "lightweight sandals for the beach in the summer."
[1365] Step 2:
[1366] The device captures the user's input text in real time and makes an API request to send that data to the server.
[1367] Step 3:
[1368] The server receives the user's input text sent from the device via the API.
[1369] Step 4:
[1370] The server passes the received text data to the emotion engine to recognize the user's emotion.
[1371] Step 5:
[1372] The emotion engine analyzes the text data and identifies the user's emotion, for example, recognizing the emotion that indicates "fun."
[1373] Step 6:
[1374] The server passes the recognized emotion data to the AI model and makes an image generation request.
[1375] Step 7:
[1376] Based on the input text and emotional data, the AI model generates an image of "summer beach sandals" that reflects "fun."
[1377] Step 8:
[1378] The server receives the images generated by the AI model and stores them in a database.
[1379] Step 9:
[1380] The image stored by the server is sent to an image search engine to search for related products and information.
[1381] Step 10:
[1382] The server receives the search results returned by the image search engine and formats them into a user-friendly format.
[1383] Step 11:
[1384] The server returns the formatted search results as an API response to be sent to the terminal.
[1385] Step 12:
[1386] The device receives the response from the server and displays the search results to the user, who can then check the results to find the product or information they are looking for.
[1387] Step 13:
[1388] If the user is not satisfied with the search results, he or she can input additional customization requests. For example, if the user desires "sandals in brighter colors," the request is input from the terminal and transmitted to the server.
[1389] Step 14:
[1390] The device obtains the user's customization requests and makes an API request to send them back to the server.
[1391] Step 15:
[1392] The server receives the customization request and passes it back to the emotion engine to reconfirm the user's emotion.
[1393] Step 16:
[1394] The emotion engine analyzes the text data included in the customization request and recognizes the user's emotion.
[1395] Step 17:
[1396] The server again passes the recognized emotion data to the AI model and makes a customized image generation request.
[1397] Step 18:
[1398] The AI model generates new images based on customized criteria and emotional data.
[1399] Step 19:
[1400] The server receives the new image and sends it back to the image search engine.
[1401] Step 20:
[1402] The server receives the search results from the image search engine again, formats them, and sends them back to the user's terminal.
[1403] Step 21:
[1404] The terminal receives the customized search results from the server and displays them to the user.
[1405] Step 22:
[1406] If the product desired by the user is not available on the market, the terminal notifies the server of this fact.
[1407] Step 23:
[1408] The server sends the user's request and the generated image to a company, which considers commercializing the image.
[1409] Step 24:
[1410] If the user has difficulty solidifying a specific image, they can access the AI concierge function from their device. When the user enters their inquiry details, the device sends the data to the server.
[1411] Step 25:
[1412] The server passes the consultation content and emotional data to the AI concierge and requests a proposal.
[1413] Step 26:
[1414] The AI concierge generates optimal suggestions based on input text and sentiment data.
[1415] Step 27:
[1416] The server receives the AI concierge's suggestions and sends them to the user's device.
[1417] Step 28:
[1418] The terminal displays the suggestions to the user, who can then make an appropriate selection.
[1419] Example 2
[1420] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1421] In the past, it was difficult for users to find a specific product that matches their mental image, and there were no systems that provided products or information that took the user's emotions into consideration. This made it difficult to provide products that matched the user's needs and emotions, and improving user satisfaction was a challenge.
[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1423] In this invention, the server includes means for a user to input the image in their mind in language, means for receiving the input and analyzing the language data using an emotion recognition engine to recognize the user's emotion, means for the server to generate an image from the language data using a generative AI model based on the recognized emotion data, means for the server to use the generated image to search for related products and information using an image search engine, and means for a terminal to display the search results to the user. This makes it possible to quickly and accurately provide products and information based on the user's needs and emotions.
[1424] "User" refers to an individual or group that uses the system to materialize the images and desires in their mind.
[1425] "Terminal" refers to an input device or display device used by a user, including a personal computer, smartphone, tablet, etc.
[1426] "Server" refers to the computer system responsible for data processing, data storage, and execution of various engines and models for the entire system.
[1427] "Language data" refers to text information entered by the user, and is made up of sentences and words that express the user's mental images and wishes.
[1428] An "emotion recognition engine" refers to an algorithm or program that analyzes the language data entered by a user and identifies the user's emotions from it.
[1429] A "generative AI model" refers to an artificial intelligence model that generates output such as images or text based on input data, and performs creative generation in response to specific prompts.
[1430] A "prompt sentence" is text data input into a generative AI model, and refers to a sentence that contains instructions for generating the output image or information.
[1431] "Image" refers to a visual representation generated by a generative AI model based on the user's language and emotional data.
[1432] An "image search engine" refers to a system or program that searches the Internet for related products and information based on an input image and returns the results.
[1433] "Search results" refers to the list of related products and information returned by an image search engine, including relevant data that may be useful to the user.
[1434] "Customization request" refers to content that the user wishes to add or change based on the search results, and includes new input for regenerating or searching.
[1435] "Company" refers to a corporation or organization that considers commercialization in response to user requests.
[1436] This invention is a system that allows users to concretely express the images in their minds and provides appropriate information and products by taking into account the user's emotions. This system is composed of a user terminal, a server, an emotion recognition engine, a generative AI model, and an image search engine.
[1437] User Input Phase
[1438] The user uses the device to input the image in their mind in words. For example, they might input "lightweight sandals to wear on the beach in the summer." The device captures the text data entered by the user in real time and sends it to a server. The device can be a PC, smartphone, tablet, or other device.
[1439] Emotion Recognition Phase
[1440] The server receives the text data sent from the device and passes it to an emotion recognition engine. The emotion recognition engine (such as "EmotionAPI") analyzes the text data and recognizes the user's emotions. For example, if the user imagines a "fun beach vacation," the emotion recognition engine will identify the emotion "enjoyment."
[1441] Image generation phase
[1442] The server requests the generative AI model to generate an image based on the emotion data obtained from the emotion recognition engine. The generative AI model (e.g., "DALL-E") generates an image based on the input text and emotion data.
[1443] Examples of prompts include:
[1444] "Lightweight sandals perfect for summer beach days. Brightly colored designs for a fun beach vacation."
[1445] Based on this, the generative AI model generates a brightly colored image of "summer flip-flops," which is then stored on a server.
[1446] Image search phase
[1447] The generated image is sent from the server to an image search engine. The image search engine (for example, Google Image Search API) searches for related products and information. The server receives the search results from the image search engine and formats them in a user-friendly format.
[1448] Result display phase
[1449] The server sends the formatted search results to the user's device, which then displays them to the user, who can then review them to find the products or information they want.
[1450] Customization Phase (Optional)
[1451] If a user is not satisfied with the search results, they can input additional customization requests. For example, if they want "sandals in a brighter color," they input that request from their device and send it to the server. The server receives the customization request and again generates a new image using the emotion recognition engine and generative AI model. The generated image is then sent to the image search engine again, and the results are provided to the user.
[1452] Commercialization request phase (optional)
[1453] If the product desired by the user does not exist on the market, the server will link the request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1454] AI Concierge Phase (Optional)
[1455] If a user has difficulty solidifying a specific image, they can call up the AI concierge function from their device. For example, if they are unsure what to wear to a colleague's wedding, they can input their concerns into their device. The server passes the concerns and emotional data to the AI concierge, which then makes optimal suggestions. For example, if the user is feeling "joy" or "elegance," the AI concierge will use that information to suggest the most appropriate outfit. These suggestions are displayed on the user's device, allowing them to make an appropriate choice.
[1456] As a result, the present invention can provide products and information that are in line with the user's feelings and desires more accurately, thereby improving user satisfaction.
[1457] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1458] Step 1: User Input Phase
[1459] The user inputs the image in their mind using language. For example, they might input "lightweight sandals to wear on the beach in the summer."
[1460] The input language data is acquired in real time by the terminal, and the terminal transmits the acquired text data to the server.
[1461] Input: User text input (e.g., "Lightweight sandals for summer beach wear")
[1462] Output: Send text data from the terminal to the server
[1463] Step 2: Emotion Recognition Phase
[1464] The server passes the text data received from the device to the emotion recognition engine, which analyzes the text data and recognizes the user's emotions.
[1465] The text "Lightweight sandals to wear on the beach in summer" is passed to an emotion recognition engine, which analyzes it and extracts the emotion data "fun."
[1466] Input: Text data received by the server
[1467] Output: Emotion data extracted by the emotion recognition engine (e.g., "enjoyment")
[1468] Step 3: Image generation phase
[1469] The server sends a prompt to the generative AI model based on the emotion data obtained from the emotion recognition engine, requesting image generation. The generative AI model generates an image based on the input text and emotion data.
[1470] Prompt: "Lightweight sandals for summer beach wear. Brightly colored designs for a fun beach vacation."
[1471] Based on this prompt, the generative AI model generates a brightly colored image of "summer flip-flops" and stores the image on a server.
[1472] Input: Prompt text and emotion data sent by the server
[1473] Output: Generated image
[1474] Step 4: Image Search Phase
[1475] The generated image is sent from the server to an image search engine, which searches for related products and information and returns the search results to the server.
[1476] The server formats the received search results in a user-friendly format.
[1477] Input: Image sent by the server
[1478] Output: Related product information returned by an image search engine
[1479] Step 5: Result display phase
[1480] The server sends the formatted search results to the user's terminal, which then displays the received search results to the user.
[1481] The user checks and selects the products and information they want from the displayed search results.
[1482] Input: The formatted search results sent by the server
[1483] Output: Search results displayed on your terminal
[1484] Step 6: Customization Phase (Optional)
[1485] If the user is not satisfied with the search results, they can input additional customization requests, such as "more vibrantly colored sandals," by entering the request on their device and sending it to the server.
[1486] The server receives the customization request, generates a new image using the emotion recognition engine and generative AI model, and then sends the generated image to the image search engine again, providing the search results to the user.
[1487] Input: Customization requests entered by the user
[1488] Output: Regenerated image and new search results
[1489] Step 7: Commercialization Request Phase (Optional)
[1490] If the product desired by the user does not exist on the market, the server will link this request to the relevant companies. The server will then send the user's request and the generated image to the companies, who will consider commercializing it.
[1491] Input: Product development requests entered by the user
[1492] Output: Requests and images sent to the company
[1493] Step 8: AI Concierge Phase (Optional)
[1494] If the user has difficulty solidifying a specific image, they can call up the AI concierge function on their device and input their inquiry. For example, if they are unsure what to wear to a colleague's wedding, they can input their inquiry into their device.
[1495] The server acquires the consultation details and emotional data and passes them to the AI concierge. The AI concierge then makes optimal suggestions based on this information. For example, if the user is feeling "joy" or "elegance," it will suggest the perfect outfit. The suggestions are displayed on the user's device, allowing the user to make an appropriate choice based on this information.
[1496] Input: Consultation content entered by the user
[1497] Output: Suggestions from the AI concierge
[1498] (Application example 2)
[1499] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1500] Conventional systems make it difficult for users to search for products based on specific product images or emotions, making it difficult to find products that fully meet the user's needs. Furthermore, with conventional text searches that are not based on emotions, there can be a gap between the products the user is looking for and the search results. The present invention aims to solve these problems by providing a more accurate product search system based on user emotions and text.
[1501] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input the image in their mind in language, means for recognizing the user's emotion from the language data using an emotion engine, means for generating an image based on the language data and emotion data using an AI model, means for the server to perform an image search using the generated image to search for related products and information, and means for the terminal to display the search results to the user. This makes it possible to quickly provide products and information based on the user's emotions and specific needs.
[1502] A "user" is an individual or entity that uses the system to accomplish a particular task.
[1503] A "mental image" is a visual representation of a particular scene or object that a user has in their mind.
[1504] "Language data" is text data that a user inputs to express an image in their mind.
[1505] A "server" is a computer system on a network that receives and processes data sent from user terminals.
[1506] An "emotion engine" is a software system that analyzes a user's language data and recognizes the emotions contained therein.
[1507] "Emotion data" is emotion information extracted from the user's language data analyzed by the emotion engine.
[1508] An "AI model" is a mathematical model that uses artificial intelligence technology to perform specific tasks.
[1509] An "image" is a visual image generated by an AI model based on a user's language and emotional data.
[1510] "Image search" is the process of searching the Internet for related products and information based on a generated image.
[1511] "Related products and information" refers to items and data found through image search based on the user's needs and emotions.
[1512] "Search results" are lists of related products and information obtained through image search.
[1513] A "terminal" is a device that can be directly operated by a user, such as a smartphone or computer.
[1514] A "customization request" is a request for improvement that a user inputs when the user is dissatisfied with the search results.
[1515] A "customized image" is an image generated again by an AI model based on customization requests.
[1516] "Interested companies" are organizations that have the potential to respond to user requests and develop or improve products.
[1517] "Marketing information" is information that includes user emotion data that companies can use as a reference when considering commercialization.
[1518] To implement this invention, the user must first input the image in their mind as concrete text. When the user inputs the text using a device such as a smartphone or computer, the data is sent to a server.
[1519] The server uses an emotion engine, such as Hugging Face's "sentiment-analysis" transformer model, to analyze the user's emotions from the received text data. The emotion engine analyzes the words and expressions in the text data and can recognize the emotions conveyed in the user's input in real time.
[1520] The server then combines the emotion data obtained from the emotion engine with the text data and generates an image using a generative AI model such as DALL-E. This is the process of generating a visual image based on the text entered by the user, further reflecting the emotion contained in the text. The generated image is then stored on the server.
[1521] The generated images are used to search for related products and information using image search engines such as the Google Image Search API. The server receives the search results returned by the image search engine and formats them in a user-friendly format.
[1522] The server sends the formatted search results to the user's device, which then displays them to the user. The user can then review the products and information displayed in the search results and make an appropriate selection.
[1523] If the user is not satisfied with the search results, they can input additional customization requests from their device. The server receives these customization requests and generates customized images using the emotion engine and generative AI model again. Further image searches are performed using these customized images, and the results are displayed to the user.
[1524] If the product desired by the user is not available on the market, the server will share the request with relevant companies. At this time, the generated image and emotion data will also be provided to the companies. Based on this, the companies can consider commercializing the product and analyze marketing information.
[1525] As a concrete example, consider a user searching for "swimsuits for a fun weekend at the beach." When the user enters this text, the emotion engine recognizes the word "fun," and the generative AI model generates a bright and cheerful image of the swimsuit. The server then searches for related products based on this image and suggests them to the user.
[1526] Example prompt sentence:
[1527] Your goal is to implement an application that recognizes emotions from the input text, such as "I'm looking for a swimsuit for a fun weekend at the beach," generates relevant images based on the emotion, and searches for and suggests related products.
[1528] In this way, the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[1529] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1530] Step 1:
[1531] The user inputs the image in their mind using language. For example, if they are looking for a swimsuit for a fun weekend at the beach, they can input this text using a device such as a smartphone or computer. The input data (text) is acquired in real time and sent to the server.
[1532] Step 2:
[1533] The server receives the text data sent from the device. Then, it uses an emotion engine (for example, Hugging Face's "sentiment-analysis" Transformer model) to analyze the user's emotions from the received text data. In this case, the input is the user's text data, and the output is emotional data such as "enjoyment."
[1534] Step 3:
[1535] The server requests a generative AI model (such as DALL-E) to generate an image based on the emotion data and text data from the emotion engine. The input is text data and emotion data, and the output is an image that reflects the emotion. Specifically, the server generates an image of a bright and cheerful beach swimsuit that reflects "fun."
[1536] Step 4:
[1537] Using the generated image, the server uses an image search engine (such as the Google Image Search API) to search for related products and information. In this case, the input is the image, and the output is a list of related products and information.
[1538] Step 5:
[1539] The server receives the search results returned by the image search engine and formats them in a user-friendly format. This formatting process ensures that the search results are displayed in a format that is appealing to the user. The input to this process is the unformatted search result data returned by the search engine, and the output is the formatted search result data.
[1540] Step 6:
[1541] The server sends the formatted search results to the user's device, which then displays the received search results to the user. The input here is the formatted search result data, and the output is the visual search results displayed on the user's device screen.
[1542] Step 7:
[1543] If the user is not satisfied with a particular product or information, they can input additional customization requests from their terminal. These requests are sent back to the server and used for the next process. The input is the text data of the user's customization requests, and the output is the customization requests sent to the server.
[1544] Step 8:
[1545] The server receives the customization request and generates a customized image using the emotion engine and generative AI model. Here, the input is the text data and emotion data of the customization request, and the output is a customized image. For example, if a user requests "sandals with a more vibrant color," the server generates an image of more vibrant sandals.
[1546] Step 9:
[1547] The server then performs another image search using the customized image, where the input is the customized image and the output is a new list of related products and information, formats the search results, and presents them to the user again.
[1548] Step 10:
[1549] If the product desired by the user does not exist on the market, the server will share this request with relevant companies. The companies will then consider commercializing the product. The input here is the user's request, the generated image, and emotion data, and the output is marketing information provided to the company and the results of consideration for commercialization.
[1550] In this way, the system of the present invention can quickly provide optimal products and information based on the user's emotions and specific needs.
[1551] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1552] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1553] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1554] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1555] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1556] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1557] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1558] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1559] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1560] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1561] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1562] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1563] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1564] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1565] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1566] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1567] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1568] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1569] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1570] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1571] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1572] The following is further disclosed regarding the above embodiment.
[1573] (Claim 1)
[1574] A means for users to input the images in their minds in words;
[1575] a server receiving the input and generating an image based on the language data using an AI model;
[1576] a means for the server to perform an image search using the generated image to search for related products and information;
[1577] means for displaying the search results to a user in a terminal;
[1578] A system including:
[1579] (Claim 2)
[1580] A means for a user to input additional customization requests based on the image search results;
[1581] A server receives the customization request and generates a customized image using the AI model again;
[1582] a means for the server to perform image search again using the customized image;
[1583] means for the terminal to display the customized search results to the user;
[1584] The system of claim 1 further comprising:
[1585] (Claim 3)
[1586] If the product desired by the user does not exist, the server will link the request to interested companies;
[1587] A means for companies to consider commercialization based on the requests;
[1588] The system of claim 1 further comprising:
[1589] (Claim 4)
[1590] When the user is unable to solidify the image in their mind, the device will have a way to access an AI concierge,
[1591] A means for the server to use the AI concierge to suggest an image based on a user's consultation;
[1592] means for the terminal to display the proposed image to the user;
[1593] The system of claim 1 further comprising:
[1594] "Example 1"
[1595] (Claim 1)
[1596] A means for users to input the images in their minds in words;
[1597] a server receiving the input and generating an image based on the language data using a generative AI model;
[1598] a means for the server to perform an image search using the generated image to search for related products and information;
[1599] means for displaying the search results to a user in a terminal;
[1600] When users have difficulty solidifying a specific image, they can call up the AI concierge function from their device.
[1601] A means for the server to pass the consultation details to the AI concierge function and propose the optimal image based on the user's wishes;
[1602] means for the terminal to display said suggestions to the user;
[1603] A system including:
[1604] (Claim 2)
[1605] A means for a user to input additional customization requests based on the image search results;
[1606] A server receives the customization request and generates a customized image using the generation AI model again;
[1607] a means for the server to perform image search again using the customized image;
[1608] means for the terminal to display the customized search results to the user;
[1609] The system of claim 1 further comprising:
[1610] (Claim 3)
[1611] If the product desired by the user does not exist, the server will link the request to interested companies;
[1612] A means for the company to consider commercialization based on said request;
[1613] The system of claim 1 further comprising:
[1614] "Application Example 1"
[1615] (Claim 1)
[1616] A means for users to input the images in their minds in words;
[1617] a server receiving the input and generating an image based on the language data using a generative AI model;
[1618] a means for the server to perform an image search using the generated image to search for related products and information;
[1619] a terminal for displaying the search results to a user, and a means for the user to input additional customization requests from the terminal;
[1620] A means for the server to receive the customization request and generate a customized image using the generation AI model again;
[1621] a means for performing an image search using the customized image again;
[1622] means for the terminal to again display the customized search results to the user;
[1623] ...
[1624] A system including:
[1625] (Claim 2)
[1626] If the product desired by the user does not exist, the server will link the request to interested companies;
[1627] A means for companies to consider commercialization based on the requests;
[1628] 10. The system of claim 1, comprising:
[1629] (Claim 3)
[1630] If the user has difficulty forming a specific image in their mind, the server will use an AI concierge function to suggest the most suitable image.
[1631] means for the terminal to display the proposed image to the user;
[1632] 10. The system of claim 1, comprising:
[1633] "Example 2: Combining Emotion Engines"
[1634] (Claim 1)
[1635] A means for users to input the images in their minds in words;
[1636] a server receiving the input and analyzing the language data using an emotion recognition engine to recognize the user's emotion;
[1637] a means for generating an image from the language data by using a generation AI model based on the recognized emotion data by the server;
[1638] A server uses the generated image to search for related products and information using an image search engine;
[1639] means for displaying the search results to a user in a terminal;
[1640] A system including:
[1641] (Claim 2)
[1642] A means for a user to input additional customization requests based on the image search results;
[1643] A server receives the customization request and generates a customized image again using an emotion recognition engine and a generation AI model;
[1644] a means for the server to perform image search again using the customized image;
[1645] means for the terminal to display the customized search results to the user;
[1646] The system of claim 1 further comprising:
[1647] (Claim 3)
[1648] A means for the user to input a description of the product they desire;
[1649] means for the server to communicate the desired content and the generated image to a company if the desired product does not currently exist;
[1650] A means for companies to consider commercialization based on the requests;
[1651] The system of claim 1 further comprising:
[1652] "Application example 2 when combining emotion engines"
[1653] (Claim 1)
[1654] A means for users to input the images in their minds in words;
[1655] a server receiving the input and using an emotion engine to recognize the user's emotion from the language data;
[1656] means for generating an image based on the language data and emotion data using an AI model;
[1657] a means for the server to perform an image search using the generated image to search for related products and information;
[1658] means for displaying the search results to a user in a terminal;
[1659] A system including:
[1660] (Claim 2)
[1661] A means for a user to input additional customization requests based on the image search results;
[1662] A server receives the customization request and generates a customized image using the AI model again;
[1663] a means for the server to perform image search again using the customized image;
[1664] means for the terminal to display the customized search results to the user;
[1665] A means for the server to store information combining the additional customization request and the generated image by the AI model;
[1666] The system of claim 1 further comprising:
[1667] (Claim 3)
[1668] If the product desired by the user does not exist, the server will link the request to interested companies;
[1669] A means for companies to consider commercialization based on the requests;
[1670] A means of analyzing user emotion data and providing it to companies as product marketing information;
[1671] The system of claim 1 further comprising: [Explanation of symbols]
[1672] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for users to input the images in their minds in words; a server receiving the input and generating an image based on the language data using an AI model; a means for the server to perform an image search using the generated image to search for related products and information; means for displaying the search results to a user in a terminal; A system including:
2. A means for a user to input additional customization requests based on the image search results; A server receives the customization request and generates a customized image using the AI model again; a means for the server to perform image search again using the customized image; means for the terminal to display the customized search results to the user; The system of claim 1 further comprising:
3. If the product desired by the user does not exist, the server will link the request to interested companies; A means for companies to consider commercialization based on the requests; The system of claim 1 further comprising:
4. When the user is unable to solidify the image in their mind, the device will have a way to access an AI concierge, A means for the server to use the AI concierge to suggest an image based on a user's consultation; means for the terminal to display the proposed image to the user; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A