system
The system addresses the lack of integrated platforms for generative AI image generation and product information by enabling seamless image output and information provision, improving user experience through efficient data transmission and reception.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Current systems lack an integrated platform for generating images using generative AI, printing them, displaying on a TV monitor, and providing information about products equipped with AI technology, requiring multiple operations and lacking timely information provision.
A system that includes means for receiving generation instructions, generating images using a generative AI model, displaying on a user terminal, transmitting to a printing or display device, and providing product information, ensuring efficient and accurate data transmission and reception.
Enables seamless integration from image generation to product information acquisition, enhancing user experience by simplifying operations and providing timely information about AI-equipped products.
Smart Images

Figure 2026062259000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, as the technology of generative artificial intelligence (AI) has evolved, the need for users to use generative AI to generate images and utilize those images in various ways has been increasing. However, current systems lack an integrated platform that can print images generated using generative AI, display them on a TV monitor, or even provide information about products utilizing generative AI technology simultaneously. As a result, users need multiple operations and procedures, lacking convenience. Also, there is a problem that relevant information cannot be provided in a timely manner to users interested in products equipped with generative AI technology.
Means for Solving the Problems
[0005] To address the aforementioned challenges, we propose a system that includes means for receiving generation instructions, means for generating images using a generation artificial intelligence model, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device for printing, means for transmitting the generated image data to a display device for display, and means for displaying information about the product equipped with artificial intelligence technology after image generation is complete. This system allows users to seamlessly perform a series of operations, from image generation by the generation AI to its output method and acquisition of product information. Furthermore, by including means for encoding and decoding the generated image data, efficiency and accuracy in the data transmission and reception process can be ensured. This system also includes a function to generate images based on text input, making it intuitive for users. With this configuration, it is possible to consistently improve the user experience of generation AI.
[0006] A "means for receiving generation instructions" refers to an interface or device that receives specific instructions regarding image generation from the user as input.
[0007] "Means using a generative artificial intelligence model" refers to computer programs or algorithms that use artificial intelligence technology to generate images based on received generation instructions.
[0008] "Means of displaying on the user terminal" refers to devices or software that display the generated image in a way that the user can see.
[0009] "Means for transmitting generated image data to a printing device and performing printing" refers to the functions and processes for transmitting generated image data to a printing device and actually performing printing.
[0010] "Means for transmitting generated image data to a display device and displaying it" refers to the function and process for transmitting generated image data to a display device such as a television and displaying the image on that device.
[0011] "Means for displaying information about products equipped with artificial intelligence technology" refers to an interface or software that, after image generation is complete, displays detailed information about products equipped with artificial intelligence technology recommended by the system to the user.
[0012] "Generating images based on text input" refers to the process where artificial intelligence generates corresponding images based on text information entered by the user.
[0013] "Means for encoding and decoding" refers to functions and processes that perform the necessary data transformations to efficiently and accurately store, transmit, and display generated image data. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined..
Mode for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0016] First, the language used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. In the following description, the program processing at each step of the system will be explained in natural language, with specific examples.
[0036] Basic configuration
[0037] 1. Preparing the user interface
[0038] Server: Starts a web server and hosts the web application that users can access. This allows users to access a form to enter instructions for image generation.
[0039] 2. Receive instructions for image generation.
[0040] User: Access the web application and enter the content of the image you want to generate as text. Submit this instruction by clicking the "Generate" button.
[0041] Terminal: Receives the entered text data and sends it to the server.
[0042] 3. Image generation using AI
[0043] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "sunset landscape" would be an instruction to generate an image depicting an evening sky and horizon.
[0044] Server: The generative artificial intelligence model generates images and sends the generation results back to the server.
[0045] 4. Displaying the generated image
[0046] Server: Sends the received image data back to the user's terminal.
[0047] Terminal: Receives image data and displays it to the user on the web application.
[0048] 5. Prepare and execute the print job.
[0049] User: Review the generated image and click the "Print" button on the web application.
[0050] Terminal: Triggers a print action and sends image data to the server.
[0051] Server: Sends the received image data to the printer's API and instructs it to print.
[0052] Printer: Print the image according to the instructions. The user can verify that the generated image is printed in high quality.
[0053] 6. Preparing and executing the TV display.
[0054] User: Review the generated image and click the "Display on TV" button.
[0055] Terminal: Triggers a display action and sends image data to the server.
[0056] Server: Sends the received image data to the TV's API and instructs it to display the image.
[0057] Television: Displays images in high resolution. Users can see that the generated images are displayed on a large screen.
[0058] 7. Display of information on AI-equipped products
[0059] Server: After image generation, printing, and display operations are complete, the server provides the user with a web application page displaying detailed information about the product, which incorporates artificial intelligence technology.
[0060] Terminal: Users can view product information they are interested in and perform actions to obtain more detailed information.
[0061] 8. Considering purchasing a Pixel product
[0062] User: View the provided product information, obtain additional information as needed, and consider purchasing the product.
[0063] Server: Provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0064] Specific example
[0065] Example 1: Creating and printing landscape photographs
[0066] User: Enters "I want a landscape photo" into the web application and clicks the "Generate" button.
[0067] Server: Sends user instructions to the generation AI and receives the generated landscape photo data.
[0068] Terminal: Displays received landscape photos.
[0069] User: After reviewing the image, click the "Print" button.
[0070] Server: Sends image data to the printer and starts printing.
[0071] Printer: Prints high-quality landscape photos.
[0072] Example 2: Generating a picture of a cat and displaying it on a TV.
[0073] User: Type "I want to generate a picture of a cat" and click the "Generate" button.
[0074] Server: Sends user instructions to the generation AI and receives the generated cat image data.
[0075] Terminal: Displays the received picture of a cat.
[0076] User: After reviewing the image, click the "Display on TV" button.
[0077] Server: Sends image data to the television and starts displaying it.
[0078] Television: Displays a picture of a cat in high resolution.
[0079] Example 3: Information acquisition for AI-powered products
[0080] User: After completing the image generation experience, view the displayed information page for AI-powered products.
[0081] Server: Provides detailed product information and answers questions via chatbot.
[0082] User: Obtain information to consider purchasing and inquire for further details as needed.
[0083] As described above, this system utilizes generative AI to generate images and provides a series of processes for outputting those images in various ways. Furthermore, by providing users with information about AI-powered products through the generation experience, it becomes possible to create a new purchasing experience.
[0084] The following describes the processing flow.
[0085] Step 1:
[0086] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[0087] Step 2:
[0088] The user accesses the web application and enters the content of the image they want to generate as text. Once the input is complete, they click the "Generate" button.
[0089] Step 3:
[0090] The terminal receives user input and sends text data to the server. The transmitted data is in JSON format.
[0091] Step 4:
[0092] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, it passes the text "sunset landscape" to the "generative AI".
[0093] Step 5:
[0094] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user terminal. The image data is usually encoded in a format such as Base64.
[0095] Step 6:
[0096] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[0097] Step 7:
[0098] After the user reviews the generated image, they click the "Print" button on the web application.
[0099] Step 8:
[0100] The terminal triggers a print action and resends the image data to the server.
[0101] Step 9:
[0102] The server receives the image data and sends it to the printer's API to issue a print command.
[0103] Step 10:
[0104] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[0105] Step 11:
[0106] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[0107] Step 12:
[0108] The device triggers a display action, and the image data is sent to the server again.
[0109] Step 13:
[0110] The server sends image data to the TV's API and issues instructions for display.
[0111] Step 14:
[0112] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[0113] Step 15:
[0114] After the server has finished generating, printing, and displaying images, it provides users with a web application containing detailed product information pages that incorporate AI technology.
[0115] Step 16:
[0116] Users can view product information and perform actions to request additional information as needed.
[0117] Step 17:
[0118] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0119] (Example 1)
[0120] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0121] Traditional image generation systems have been criticized for their complex user experience, as the process of generating, displaying, and printing images is cumbersome. Furthermore, they lack mechanisms for quickly providing product information related to the generated images, missing opportunities to increase user purchasing intent. Additionally, there are insufficient methods for users to easily ask questions about products and receive answers.
[0122] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0123] In this invention, the server includes means for hosting a web application accessible by a user terminal, means for receiving generation instructions, means for using a generative artificial intelligence model that generates images based on the generation instructions, means for returning and displaying the generated image data to the user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for displaying information about a product equipped with artificial intelligence technology after the image generation is complete, and means for the user to ask questions about the product through a chatbot and receive responses. This allows the user to easily and quickly perform a series of processes from image generation to display, printing, and product information acquisition, resulting in a high-quality user experience.
[0124] A "user terminal" is a device that a user directly operates to access web applications and the internet.
[0125] A "web application" is a software application that operates via the internet or an intranet and is accessible to users through a web browser.
[0126] "Generation instructions" refer to the text or commands that the user enters to generate an image.
[0127] A "generative artificial intelligence model" is a system that includes machine learning algorithms for generating content such as images based on specified input data.
[0128] A "server" is a computer system on a network that hosts web applications and processes requests from users.
[0129] "Image data" refers to the digital data format of a generated image, which is data that can be displayed and printed.
[0130] A "printing device" is a hardware device that receives digital image data and prints it onto physical paper.
[0131] A "display device" is a device that receives image data and displays it on a screen in high resolution.
[0132] "Product information" refers to data including specifications, reviews, and purchase information related to products equipped with artificial intelligence technology.
[0133] A "chatbot" is conversational software that automatically responds to questions from users.
[0134] Basic configuration
[0135] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. This system includes the following elements:
[0136] 1. Preparing the user interface
[0137] Server: Start a web server (e.g., Apache®, Nginx) and host a web application using Python and a framework like Django or Flask. This allows users to access a webpage with a form to input instructions for image generation.
[0138] 2. Receive instructions for image generation.
[0139] User: Access the web application and enter the content of the image you want to generate as text. For example, use a prompt like "Sunset Landscape". This prompt is submitted when you click the "Generate" button.
[0140] Terminal: Captures the entered text data and sends it to the server.
[0141] 3. Image generation using AI
[0142] Server: Analyzes the received text data and sends it as a prompt to a generative AI model (e.g., OpenAI®'s DALL-E or GPT-4®). Examples of prompts include "sunset landscape" and "picture of a cat."
[0143] Server: Receives image data generated by the generation AI model based on instructions and verifies its format (e.g., PNG, JPEG).
[0144] 4. Displaying the generated image
[0145] Server: Sends the generated image data back to the user's terminal.
[0146] Terminal: Renders received image data on a web browser and displays it to the user.
[0147] 5. Prepare and execute the print job.
[0148] User: Review the displayed image and click the "Print" button on the web application.
[0149] Terminal: Triggers a print action and sends image data to the server.
[0150] Server: Sends image data to the printer's API (e.g., HP's Printer API) and issues a print command.
[0151] Printer: Prints high-quality images using the received image data.
[0152] 6. Preparing and executing the TV display.
[0153] User: Check the displayed image and click the "Display on TV" button.
[0154] Terminal: Triggers a display action and sends image data to the server.
[0155] Server: Sends image data to the TV's API (e.g., Chromecast API) and instructs it to display the image.
[0156] Television: Displays images in high resolution.
[0157] 7. Display of information on AI-equipped products
[0158] Server: After the user completes operations such as image generation, printing, and display, the server provides a page in the web application that displays information about the product, which incorporates artificial intelligence technology. The information page includes detailed product specifications, user reviews, and purchase information.
[0159] Terminal: Displays product information on a webpage, allowing users to view details.
[0160] 8. Considering purchasing a Pixel product
[0161] User: Browse the displayed AI product information page and review the details. Enter questions into the chatbot as needed.
[0162] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[0163] Specific examples of operation
[0164] Example 1: Creating and printing landscape photographs
[0165] User: Enters "I want a landscape photo" into the web application and clicks the generate button.
[0166] Server: Sends instructions as prompts to the AI model generating the data (e.g., DALL-E) and receives the generated landscape photos.
[0167] Terminal: The landscape photo is displayed in a web browser, and the user clicks the print button after confirming it.
[0168] Server: Sends image data to the printer's API and issues a print command.
[0169] Printer: Prints landscape photos in high quality.
[0170] Example 2: Generating a picture of a cat and displaying it on a TV.
[0171] User: Enters "I want to generate a picture of a cat" into the web application and clicks the generate button.
[0172] Server: Sends instructions as prompts to the AI model that generates the data, and receives the generated picture of a cat.
[0173] Device: A picture of a cat is displayed on the webpage, and the user clicks the "Show on TV" button after confirming it.
[0174] Server: Sends image data to the TV's API and issues display instructions.
[0175] Television: Displays a picture of a cat in high resolution.
[0176] Example 3: Information acquisition for AI-powered products
[0177] User: After completing the image generation experience, view the displayed AI product information page to learn more details.
[0178] Server: Provides detailed product specifications, user reviews, and purchase information, and answers questions via chatbot functionality.
[0179] User: Obtain information to help with purchase considerations and resolve questions via chatbots or web pages as needed.
[0180] This invention allows users to easily and quickly perform a series of processes from image generation to display, printing, and acquisition of product information. The system aims to provide a high-quality user experience and enhance user purchasing intent.
[0181] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0182] Program processing steps
[0183] Step 1:
[0184] User: Accesses the web application and enters instructions for image generation. Specifically, the user enters prompt text such as "Sunset Landscape" or "Picture of a Cat" into the text input field and clicks the "Generate" button.
[0185] Input: Prompt text (e.g., "Sunset scenery")
[0186] Output: Text data sent by user interaction
[0187] Step 2:
[0188] Terminal: Sends the entered text data to the server. Specifically, when the user clicks the "Generate" button, that text data is sent to the server as an HTTP request.
[0189] Input: Text data entered by the user
[0190] Output: Text data sent to the server
[0191] Step 3:
[0192] Server: Analyzes the received text data and sends it as a prompt to the generative AI model. Natural language processing (NLP) techniques are used for analysis, and API requests are generated to the generative AI model (e.g., OpenAI's DALL-E or GPT-4).
[0193] Input: Text data received from the device
[0194] Output: API request to the generated AI model
[0195] Step 4:
[0196] Generative AI model: Generates images based on text data. The generated image data is returned to the server.
[0197] Input: Prompt message sent from the server (API request)
[0198] Output: Generated image data
[0199] Step 5:
[0200] Server: Retrieves image data received from the generated AI model and sends it back to the user's terminal. Here, it checks the image data format (e.g., PNG, JPEG) and encodes it in the appropriate format.
[0201] Input: Image data returned from a generative AI model
[0202] Output: Image data to be sent to the user's terminal.
[0203] Step 6:
[0204] Terminal: Renders received image data on a web browser and displays it to the user.
[0205] Input: Image data returned from the server
[0206] Output: Image displayed in the web browser
[0207] Step 7:
[0208] User: Review the generated image and click the "Print" or "Display on TV" button on the web application.
[0209] Input: Image data displayed in a web browser
[0210] Output: User instructions for printing or displaying.
[0211] Step 8:
[0212] Terminal: Triggers user actions (print or view) and sends image data to the server.
[0213] Input: User's print or display instructions
[0214] Output: Image data sent to the server
[0215] Step 9:
[0216] Server: Sends image data to the corresponding device's API to instruct it to print or display. For printing, it sends the data to the printer's API (e.g., HP's Printer API), and for display, it sends the data to the TV's API (e.g., Chromecast API).
[0217] Input: Image data sent from the device
[0218] Output: Image data sent to a printer or television.
[0219] Step 10:
[0220] Printer or television: Receives image data and displays or prints it on the respective device. A printer prints the image on paper, while a television displays the image in high resolution.
[0221] Input: Image data sent from the server
[0222] Output: Printed image or displayed image
[0223] Step 11:
[0224] Server: After the user completes the image generation, printing, and display operations, the server provides a page in the web application that displays product information powered by artificial intelligence technology.
[0225] Input: User operation completion information
[0226] Output: Product information display page
[0227] Step 12:
[0228] User: View information about AI-powered products on a webpage and check the details. Enter questions into the chatbot as needed.
[0229] Input: Viewing product information page
[0230] Output: Question sent to the chatbot
[0231] Step 13:
[0232] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[0233] Input: Question from a user
[0234] Output: Chatbot or human response
[0235] (Application Example 1)
[0236] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0237] Conventional image generation systems had a cumbersome process for users to apply their generated images to specific products and lacked intuitive preview functionality. As a result, users could not visually confirm how the generated images would appear on the actual product, making it difficult to achieve a satisfactory purchasing experience. Furthermore, the means of transmitting the generated images to various output devices for appropriate display and printing were limited.
[0238] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0239] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for applying the generated images to a specific product, means for displaying the result of applying the generated images to a specific product as a three-dimensional preview, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This allows the user to intuitively apply the generated images to a specific product and confirm them in a three-dimensional preview. Furthermore, the generated images can be easily displayed and printed on various output devices, improving the user's purchasing experience.
[0240] A "means for receiving generation instructions" refers to a device or system that receives input from a user to instruct image generation.
[0241] "Means using a generative artificial intelligence model that generates images based on generation instructions" refers to devices or systems that utilize artificial intelligence technology to generate images based on user instructions.
[0242] "Means for displaying generated images on a user's terminal" refers to devices or systems for displaying generated image data on a user's terminal.
[0243] "Means for transmitting generated image data to a printing device and performing printing" refers to a device or system for transmitting generated image data to a printing device and printing the image using the printing device.
[0244] "Means for transmitting generated image data to a display device and displaying it" refers to a device or system for transmitting generated image data to a display device and displaying the image on the display device.
[0245] "Means for applying a generated image to a specific product" refers to a device or system for applying a generated image to a desired product.
[0246] "Means for displaying the result of applying a generated image to a specific product as a three-dimensional preview" refers to a device or system for applying a generated image to a specific product and visualizing and displaying the result in three dimensions.
[0247] "Means for displaying information about a product equipped with artificial intelligence technology after the generation of the aforementioned image is completed" refers to a device or system for providing information about a product equipped with artificial intelligence technology after the generation process is completed.
[0248] The following describes the specific system configuration and program processing for implementing this invention. In particular, it describes the detailed process of generating images using a generative AI model and applying them to a product.
[0249] Basic System Configuration
[0250] 1. Preparing the user interface
[0251] Server: Starts a web server and hosts a web application for users to access. This web application includes a form for entering instructions for image generation.
[0252] 2. Receive instructions for image generation.
[0253] User: Access the web application and enter the content of the image you want to generate as text. Submit this text instruction by clicking the "Generate" button.
[0254] 3. Image generation using AI
[0255] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "Colorful geometric patterns" becomes an instruction to generate an image based on its content.
[0256] Server: The generative artificial intelligence model generates images and sends the results back to the server.
[0257] 4. Displaying the generated image
[0258] Server: Sends the received image data back to the user's terminal.
[0259] Terminal: Receives image data and displays it to the user on the web application.
[0260] 5. Application of images to products
[0261] User: Review the generated image and select the option to apply it to a specific product (e.g., T-shirt, cup, etc.).
[0262] Server: Uses image data to place and apply images to specified products.
[0263] Terminal: Displays a 3D preview of the generated product for the user to review.
[0264] 6. Prepare and execute printing.
[0265] User: After reviewing the generated image, decide to order the product and click the "Print" button.
[0266] Server: Receives print instructions and sends the generated image data to the printing device.
[0267] Printing device: Follow the instructions and print the generated image onto the product.
[0268] 7. Preparing and executing the TV display.
[0269] User: After reviewing the generated image, click the "Display on TV" button.
[0270] Server: Triggers a display action and sends image data to the display device.
[0271] Display device: Displays images in high resolution.
[0272] 8. Display of information on AI-equipped products
[0273] Server: After image generation, printing, and display operations are complete, the server displays a page on the web application that provides users with detailed information about the product, which incorporates artificial intelligence technology.
[0274] Hardware and software to be used
[0275] hardware
[0276] User's smartphone
[0277] Server (e.g., cloud service)
[0278] Printing device
[0279] Display device (home TV)
[0280] Software
[0281] Front - end: React
[0282] Back - end: Flask (Python)
[0283] Image generation: OpenAI API
[0284] Specific example
[0285] Examples of prompt sentences
[0286] "Colorful geometric patterns"
[0287] "Scenery of the universe"
[0288] "Cute cat illustration"
[0289] For example, a user accesses a web application, enters "Colorful geometric patterns", and clicks the "Generate" button. The server that receives this instruction uses the generative AI model to generate an image based on the specified content. The generated image is first displayed on the user terminal. Next, the user applies the image to a specific product and views it in a 3D preview. If satisfied with this preview, the user can actually send the image to the printing device to print it on the product. Also, the generated image can be displayed on a home TV. Finally, detailed information about the product equipped with artificial intelligence technology is provided to the user who has completed the image generation experience.
[0290] In summary, this invention provides a system that allows users to intuitively generate images and quickly apply them to products for visualization. Furthermore, user satisfaction can be enhanced by displaying or printing the generated images on various output devices.
[0291] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0292] Step 1:
[0293] The user enters a prompt.
[0294] The user accesses the web application and enters a text prompt for image generation (e.g., "Colorful geometric pattern"). Once the input is complete, they click the "Generate" button.
[0295] Input: Text prompt (e.g., "Colorful geometric patterns")
[0296] Output: Generation instructions
[0297] Specific operation: The user enters a prompt message in the text input field and clicks a button, which sends the instruction to the server.
[0298] Step 2:
[0299] The server receives the generation instruction.
[0300] The server receives the prompt message sent by the user and parses it.
[0301] Input: Generation Instructions
[0302] Output: Parsed prompt message
[0303] Specific operation: The server receives an HTTP request and extracts the prompt text.
[0304] Step 3:
[0305] The server sends a request to the generative AI model.
[0306] Based on the parsed prompt text, the server sends an image generation request to the generative AI model.
[0307] Input: Parsed prompt text
[0308] Output: Request for the generative AI model
[0309] Specific operation: The server sends the prompt text to the API of the generative AI model in an appropriate format.
[0310] Step 4:
[0311] The generative AI model generates an image.
[0312] The generative AI model generates an image based on the prompt text and returns the result to the server.
[0313] Input: Request for the generative AI model
[0314] Output: Generated image data ]>
[0315] Specific operation: The generative AI model performs internal processing, generates an image based on the prompt text, and sends the data to the server.
[0316] Step 5:
[0317] The server receives the generated image and sends it to the user terminal. <了
[0318] The server receives the generated image data and returns it to the user terminal.
[0319] Input: Generated image data
[0320] Output: Sending image data to the user terminal
[0321] Specific operation: The server sends the received image data to the user's terminal as an HTTP response.
[0322] Step 6:
[0323] Display images on the user's terminal.
[0324] The user terminal displays the received image data on the web application.
[0325] Input: Image data sent to the user terminal
[0326] Output: Image display on a web application
[0327] Specific operation: The user's device browser renders the image data and displays it on the screen.
[0328] Step 7:
[0329] Users apply images to the product.
[0330] The user reviews the generated image and selects the option to apply it to a specific product (e.g., a T-shirt, a cup).
[0331] Input: Generated image and product selection
[0332] Output: Image applied to the product
[0333] Specific operation: The user clicks a product selection option and sends an apply request to the server.
[0334] Step 8:
[0335] The server applies the image to the product.
[0336] The server applies the generated image to a specific product and generates the corresponding data.
[0337] Input: Generated image and product selection information
[0338] Output: Image data applied to the product
[0339] Specific operation: The server places images into the specified product template and generates the applied data.
[0340] Step 9:
[0341] The server sends a 3D preview to the user terminal.
[0342] The server sends the generated product data to the user terminal in a three-dimensional preview format.
[0343] Input: Image data applied to the product
[0344] Output: Sending a 3D preview to the user's terminal
[0345] Specific operation: The server generates data for the 3D preview and sends it to the user's terminal as an HTTP response.
[0346] Step 10:
[0347] The user terminal displays a 3D preview.
[0348] The user terminal displays the received 3D preview on the web application.
[0349] Input: Data for 3D preview
[0350] Output: Three-dimensional preview display on screen
[0351] Specific operation: The user's browser renders a 3D preview and displays it on the screen.
[0352] Step 11:
[0353] The user issues a print command.
[0354] The user reviews the 3D preview, and if satisfied, clicks the "Print" button to initiate printing.
[0355] Input: User's print instructions
[0356] Output: Print request to server
[0357] Specific action: The user clicks a button and sends a print command to the server.
[0358] Step 12:
[0359] The server sends image data to the printer.
[0360] The server receives the print command and sends the generated image data to the printing device.
[0361] Input: User's print instructions and image data.
[0362] Output: Print instructions to the printer
[0363] Specific operation: The server sends image data to the printer's API and issues a print command.
[0364] Step 13:
[0365] The printing device prints the image.
[0366] The printing device prints an image onto the specified product based on the received image data.
[0367] Input: Print instructions to the printer and image data.
[0368] Output: Printed product
[0369] Specific operation: The printing device prints an image onto the product based on the data it receives.
[0370] Step 14:
[0371] The user displays an image on the display device.
[0372] The user clicks the "Display on TV" button to display the generated image on a display device such as a home television.
[0373] Input: User's display instructions
[0374] Output: Display request to the server
[0375] Specific operation: The user clicks a button and sends a display instruction to the server.
[0376] Step 15:
[0377] The server sends image data to the display device.
[0378] The server receives the display instruction and sends the generated image data to the display device.
[0379] Input: User display instructions and image data
[0380] Output: Display instructions for the display device.
[0381] Specific operation: The server sends image data to the display device's API and issues a display instruction.
[0382] Step 16:
[0383] The display device displays an image.
[0384] The display device displays the image in high resolution based on the received image data.
[0385] Input: Display instructions and image data for the display device.
[0386] Output: Displayed image
[0387] Specific operation: The display device displays an image on the screen based on the data it receives.
[0388] Step 17:
[0389] The server displays information about products equipped with AI technology.
[0390] After the server completes image generation, printing, and display operations, it displays a page on the web application that provides users with detailed information about the AI-powered product.
[0391] Input: User's operation complete
[0392] Output: Information on products equipped with AI technology
[0393] Specific action: The server updates the content of the webpage and provides the user with detailed information.
[0394] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0395] This invention relates to a generative AI system that combines an emotion engine to recognize user emotions. This system receives image generation instructions from the user, generates images using a generative artificial intelligence model, and outputs the generated images to a printer or display device. Furthermore, it displays information about products equipped with artificial intelligence technology after image generation. The system incorporates an emotion engine that recognizes the user's emotional state and adjusts generation instructions based on that state.
[0396] Basic configuration
[0397] 1. Preparing the user interface
[0398] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[0399] 2. Initiation of emotion recognition
[0400] The user accesses the web application and clicks a specific button to initiate emotion recognition.
[0401] The device collects user interaction and input data and sends it to the emotion engine.
[0402] 3. Analysis of emotional state
[0403] The server uses an emotion engine to analyze the user's emotional state. For example, it identifies emotions based on the user's input speed, input content, and facial expression data (if acquired using a camera).
[0404] 4. Receive instructions for image generation.
[0405] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[0406] The emotion engine takes the user's emotional state into consideration and makes adaptive adjustments to the generation instructions.
[0407] The terminal sends the adjusted generation instruction data to the server.
[0408] 5. Image generation using AI
[0409] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, based on the input text "Sunset Landscape," it generates an image that contains emotionally positive elements.
[0410] 6. Displaying the generated image
[0411] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user's terminal.
[0412] The device receives image data and displays it to the user on the web application. The user then views the generated image on the screen.
[0413] 7. Prepare and execute the print job.
[0414] The user reviews the generated image and clicks the "Print" button on the web application.
[0415] The terminal triggers a print action and resends the image data to the server.
[0416] The server receives the image data and sends it to the printer's API to issue a print command.
[0417] The printer prints the image in high quality according to the instructions. The user receives the printed image.
[0418] 8. Preparing and executing the TV display.
[0419] After the user reviews the generated image, they click the "Display on TV" button.
[0420] The device triggers a display action, and the image data is sent to the server again.
[0421] The server sends image data to the TV's API and issues instructions for display.
[0422] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[0423] 9. Information display for AI-equipped products
[0424] After the server has finished generating, printing, and displaying images, it provides users with detailed information about the AI-powered product via a web application.
[0425] The device displays the product information page to the user.
[0426] 10. Considering purchasing a Pixel product
[0427] Users view product information and, if interested, take actions to request additional information.
[0428] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0429] Specific example
[0430] Example 1: Emotion-driven landscape photography and printing
[0431] The user enters "I want a landscape photo" and clicks the "Generate" button.
[0432] The emotion engine analyzes user input and interaction data to identify positive emotional states.
[0433] The server sends an instruction to the AI model to generate a "positive sunset landscape."
[0434] The server receives the generated landscape photo data and sends it to the terminal.
[0435] The device displays landscape photos it has received.
[0436] After the user reviews the image, they click the "Print" button.
[0437] The server sends the image data to the printer and starts printing.
[0438] The printer prints high-quality landscape photographs.
[0439] Example 2: Generation and display of cat images based on emotions
[0440] The user enters "I want to generate a picture of a cat" and clicks the "Generate" button.
[0441] The emotion engine analyzes the user's interaction data and determines that the user is slightly tired.
[0442] The server sends an instruction to the AI model to generate a "relaxing picture of a cat."
[0443] The server receives the generated cat image data and sends it to the terminal.
[0444] The device displays the image of a cat it received.
[0445] After the user reviews the image, they click the "Display on TV" button.
[0446] The server sends the image data to the television and begins displaying it.
[0447] The TV displays a picture of a cat in high resolution.
[0448] Example 3: Information retrieval for AI-powered products based on emotions
[0449] After the user completes the image generation experience, they can view information pages about AI-powered products on the web application.
[0450] The emotion engine recognizes the user's level of excitement and prioritizes displaying highly relevant product information.
[0451] The server provides detailed product information, and a chatbot answers questions.
[0452] Users can obtain information to consider purchasing and inquire about details as needed.
[0453] In this way, a system is realized that takes user emotions into consideration and provides a consistent service from image generation using generative AI to output methods and product information acquisition. By incorporating an emotion engine, the user experience can be further personalized and satisfaction can be increased.
[0454] The following describes the processing flow.
[0455] Step 1:
[0456] The server starts up a web server and hosts a web application for users to access. This allows users to access a form to enter instructions for image generation.
[0457] Step 2:
[0458] The user accesses a web application and clicks a specific button to initiate emotion recognition. This action causes the emotion engine to begin collecting data about the user.
[0459] Step 3:
[0460] The device collects user interaction data (such as input speed, input content, and in some cases, facial recognition data from the webcam) and sends it to the emotion engine.
[0461] Step 4:
[0462] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the collected data to determine whether the user is in a positive, negative, relaxed, or other emotional state.
[0463] Step 5:
[0464] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[0465] Step 6:
[0466] The emotion engine considers the user's emotional state and makes adaptive adjustments to the "generate" instructions. For example, if the user is in a positive state, it adds instructions to include vibrant colors and bright scenes in the generated image.
[0467] Step 7:
[0468] The terminal sends the adjusted generation instruction data to the server. The transmitted data includes text as well as adjustment information for the emotion engine.
[0469] Step 8:
[0470] The server analyzes the received text data and sentiment adjustment data, and sends an image generation request to the generative artificial intelligence model. For example, when providing the generative AI model with the text "sunset landscape," instructions are also given to include positive elements.
[0471] Step 9:
[0472] The server receives the generated image data returned from the generated artificial intelligence model and sends it to the user's terminal. The image data is usually encoded in a format such as Base64.
[0473] Step 10:
[0474] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[0475] Step 11:
[0476] After the user reviews the generated image, they click the "Print" button on the web application.
[0477] Step 12:
[0478] The terminal triggers a print action and resends the image data to the server.
[0479] Step 13:
[0480] The server receives the image data and sends it to the printer's API to issue a print command.
[0481] Step 14:
[0482] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[0483] Step 15:
[0484] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[0485] Step 16:
[0486] The device triggers a display action, and the image data is sent to the server again.
[0487] Step 17:
[0488] The server sends image data to the TV's API and issues instructions for display.
[0489] Step 18:
[0490] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[0491] Step 19:
[0492] After the server completes the image generation, printing, and display operations, it provides users with detailed information about the product, which incorporates artificial intelligence technology, through a web application.
[0493] Step 20:
[0494] Users view the displayed product information and, if interested, take action to request additional information.
[0495] Step 21:
[0496] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0497] (Example 2)
[0498] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0499] Conventional image generation systems often fail to consider the user's emotional state when issuing generation instructions, resulting in generated images that do not always match the user's expectations or feelings. Furthermore, the limited output formats and methods for generated images posed a challenge, hindering user convenience.
[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0501] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for adjusting the generation instructions using an emotion engine that recognizes the user's emotional state, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This makes it possible to generate images that take the user's emotional state into consideration, thereby increasing user satisfaction. In addition, the generated images can be output to a variety of devices, improving user convenience.
[0502] "Generation instructions" are data entered by the user to specify the content and conditions of the image that should be generated.
[0503] A "generative artificial intelligence model" is an artificial intelligence algorithm or system that generates images based on instructions or data input by a user.
[0504] A "user terminal" refers to a device such as a computer or smartphone that a user operates, and which provides an interface for image generation through a web application.
[0505] A "printing device" is a hardware device used to physically print generated images, and primarily refers to a printer.
[0506] A "display device" is a hardware device used to visually present generated images to a user, and includes televisions and monitors.
[0507] An "emotion engine" is software or a system that analyzes user interaction data and input to identify emotional states and adjusts generation instructions based on those results.
[0508] "Information about products equipped with artificial intelligence technology" refers to detailed information about products incorporating artificial intelligence technology, which is provided to the user after the display of the generated image is complete.
[0509] This invention relates to an image generation system that incorporates an emotion engine to recognize user emotions. The system receives image generation instructions from the user and generates images using a generation artificial intelligence model. The generated images are then output to a printer or display device, and information about the product, which incorporates artificial intelligence technology, is displayed after image generation.
[0510] composition
[0511] The system of the present invention has the following components.
[0512] 1. Preparing the user interface
[0513] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is implemented using HTML / CSS / JavaScript (registered trademark).
[0514] 2. Initiation of emotion recognition
[0515] The user clicks the "Start Emotion Recognition" button in the web application. The device captures the click event and prepares to send interaction data to the emotion engine. At this time, it utilizes services such as the Emotion API from Microsoft® Azure®.
[0516] 3. Analysis of emotional state
[0517] The server sends the received interaction data to the emotion engine, which then analyzes the user's emotional state. For example, it analyzes input content, input speed, and facial expression data captured by the camera.
[0518] 4. Receive instructions for image generation.
[0519] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the emotional state and makes adaptive adjustments to the generation instructions. The device then sends the adjusted generation instruction data to the server.
[0520] 5. Image generation using AI
[0521] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[0522] 6. Displaying the generated image
[0523] The server receives the generated image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[0524] 7. Prepare and execute the print job.
[0525] The user reviews the generated image and clicks the "Print" button. The device resends the image data to the server. The server sends the data to the printer's API (such as Google Cloud Print) and prints the image.
[0526] 8. Preparing and executing the TV display.
[0527] After the user reviews the generated image, they click the "Display on TV" button. The device sends the image data to the server, which then sends it to the TV's API (such as GOOGLE CAST® or Samsung Smart View API) for display.
[0528] 9. Information display for AI-equipped products
[0529] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[0530] 10. Considering purchasing a Pixel product
[0531] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[0532] Specific example
[0533] Example 1: Creating and printing landscape photographs based on emotions.
[0534] User: Enter "I want a landscape photo" and click the "Generate" button.
[0535] Emotion Engine: Analyzes user input and emotional states to identify positive emotional states.
[0536] Server: Sends the instruction "positive sunset scenery" to the generating artificial intelligence model.
[0537] Server: Receives the generated image data and sends it to the terminal.
[0538] Terminal: Receives and displays image data.
[0539] User: Click the "Print" button.
[0540] Server: Sends image data to the printer API and starts printing.
[0541] Printer: Prints high-quality landscape photos.
[0542] Example 2: Generation of cat images based on emotions and display on television
[0543] User: Enter "I want to generate a picture of a cat" and click the "Generate" button.
[0544] Emotion Engine: Analyzes user interaction data to recognize fatigue levels.
[0545] Server: Instructs the artificial intelligence model to generate a "relaxing picture of a cat".
[0546] Server: Receives the generated cat image data and sends it to the terminal.
[0547] Terminal: Receives and displays image data.
[0548] User: Click the "Display on TV" button.
[0549] Server: Sends image data to the TV API and starts displaying it on the TV.
[0550] Television: Displays a picture of a cat in high resolution.
[0551] Example 3: Information acquisition for AI-powered products based on emotions
[0552] User: After completing the image generation experience, they browse the information page for AI-powered products in the web application.
[0553] Emotion Engine: Recognizes the user's emotional state and prioritizes displaying highly relevant product information.
[0554] Server: Provides detailed product information and answers questions via chatbot.
[0555] User: Obtain information to consider purchasing and inquire for further details as needed.
[0556] In this way, by incorporating an emotion engine, we can personalize the user experience and create a system that enhances satisfaction. The collaboration between the generative AI model and the emotion engine enables the generation and output of appropriate images that take into account the user's emotional state.
[0557] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0558] Step 1: Preparing the User Interface
[0559] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is constructed using HTML, CSS, and JavaScript. Users access it through a web browser and are presented with a form to input instructions for image generation.
[0560] Input: User access request
[0561] Output: Display of a form for entering image generation instructions.
[0562] Specific operation: The server starts up, and the web application waits for user requests. When a user accesses the web application, an instruction input form is displayed.
[0563] Step 2: Initiating Emotion Recognition
[0564] The user clicks the "Start emotion recognition" button in the web application. The device captures the click event and collects user interaction data (facial expressions, typing speed, text input, etc.).
[0565] Input: User click event
[0566] Output: Start of interaction data collection
[0567] Specific actions: The user clicks a button, the device detects the click event and turns on the camera. The device prepares to collect facial expression data and text input data.
[0568] Step 3: Analyze your emotional state
[0569] The server sends interaction data from the terminal to an emotion engine (for example, Microsoft Azure's Emotion API) to analyze the user's emotional state.
[0570] Input: Interaction data from the device
[0571] Output: Sentiment analysis results from the emotion engine
[0572] Specific operation: The server receives data sent from the terminal and sends it to the emotion engine. The emotion engine identifies the emotional state and returns the result to the server.
[0573] Step 4: Receive instructions for image generation.
[0574] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the user's emotional state and adaptively adjusts the generation instructions. The device then sends the adjusted generation instruction data to the server.
[0575] Input: User's image generation instructions
[0576] Output: Adjusted generation instruction data
[0577] Specific operation: The user enters the image content as text and clicks the "Generate" button. The device sends this input data to the server, which then adjusts the instructions via the emotion engine.
[0578] Step 5: Image generation using AI
[0579] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[0580] Input: Adjusted generation instruction data
[0581] Output: Generated image data
[0582] Specific operation: The server analyzes the text data and sends it to the generative AI model as a prompt. The generative AI model generates an image and sends it back to the server.
[0583] Step 6: Displaying the generated image
[0584] The server receives the generated image data returned from the generated AI model and sends it to the terminal. The terminal receives the image data and displays it to the user on the web application.
[0585] Input: Generated image data
[0586] Output: Image display on a web application
[0587] Specific operation: The server receives image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[0588] Step 7: Prepare and execute print
[0589] The user reviews the generated image and clicks the "Print" button. The device triggers the print action and resends the image data to the server. The server sends the image data to the printer API and starts printing.
[0590] Input: User's print instructions
[0591] Output: Printed image
[0592] Specific operation: The user reviews the image and clicks the "Print" button. The device sends the image data to the server, and the server sends a print command to the printer API. The printer prints the image in high quality.
[0593] Step 8: Prepare and run the TV display.
[0594] After the user reviews the generated image, they click the "Display on TV" button. The device triggers the display action and sends the image data to the server again. The server sends the image data to the TV's API and starts displaying it.
[0595] Input: User's display instructions
[0596] Output: Image displayed on the TV
[0597] Specific steps: The user views the image and clicks the "Display on TV" button. The device sends the image data to the server, and the server sends a display command to the TV API. The TV displays the image in high resolution.
[0598] Step 9: Displaying information about AI-powered products
[0599] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[0600] Input: State after operation is complete
[0601] Output: Product Information Page
[0602] Specific operation: The server adds product information to the web application, and the terminal displays the product information page to the user.
[0603] Step 10: Consider purchasing a Pixel product
[0604] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[0605] Input: User's question
[0606] Output: Automated chatbot response and transfer to a human resource representative.
[0607] Specific operation: The user views product information and enters a question into the chatbot. The server provides an automated response and, if necessary, transfers the user to a human representative.
[0608] (Application Example 2)
[0609] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0610] Conventional image generation systems do not provide personalized image generation or content recommendations based on the user's emotional state. This results in a limited user experience and potentially lower satisfaction. Furthermore, there is a need to integrate emotion recognition technology into image generation and content recommendations to provide services that are more emotionally resonant.
[0611] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0612] In this invention, the server includes means for receiving generation instructions, means for using a generative artificial intelligence model to generate images based on the generation instructions, means for analyzing the user's emotional state using emotion recognition means, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, and means for displaying information about products equipped with artificial intelligence technology after the image generation is complete. This enables personalized image generation and content recommendations according to the user's emotional state.
[0613] A "generation instruction" refers to a request regarding images or content that the user will generate.
[0614] A "generative artificial intelligence model" is a part of an artificial intelligence system that includes algorithms and databases for generating images and content based on user instructions and data.
[0615] "Emotion recognition means" refers to a part of a system that includes hardware and software for analyzing the user's emotional state from facial expressions, tone of voice, input speed, etc.
[0616] "Generated image data" refers to the digital representation of image information generated by a generative artificial intelligence model.
[0617] A "user terminal" is a device used to display or print generated images and recommended content.
[0618] A "printing device" is a machine used to print generated image data into a physical format.
[0619] A "display device" is a screen or monitor used to visually display generated image data or recommended content.
[0620] "Products equipped with artificial intelligence technology" refer to goods and services that provide functions using artificial intelligence technologies such as image generation and emotion recognition.
[0621] "Personalized content" refers to content that is customized for a specific user based on their emotional state and preferences.
[0622] The system for implementing this invention is an emotion recognition and content generation system that recognizes the user's emotions and generates and recommends personalized content accordingly.
[0623] This system includes a means of receiving generation instructions from users. A server hosts a web application with a user interface, from which users input instructions for generating images and content. These instructions include the content of the images and content desired by the user.
[0624] The server uses a generative artificial intelligence model to generate images and content based on generation instructions. This model has the ability to generate high-quality images and content based on input such as text prompts.
[0625] Next, the emotion recognition system analyzes the user's emotional state. The emotion recognition system includes a camera and microphone to capture the user's facial expressions, voice tone, input speed, etc., and software to analyze the input data. EmotionEngine, using TENSORFLOW®, identifies the emotional state.
[0626] The server displays the generated image data and content on the user's terminal. The user's terminal is a compatible device such as a smartphone, smart glasses, or head-mounted display.
[0627] Furthermore, it includes means for sending the generated image data to a printing device and printing it in physical form. To achieve this, the server uses the printer's API to send the image data and issue print commands.
[0628] In addition, the system includes means for transmitting the generated image data to a display device and displaying it on a large screen or high-resolution display.
[0629] Finally, the server displays information about the AI-powered product after image generation is complete. This allows users to obtain detailed information and purchase options for the AI-powered product associated with the generated content.
[0630] For example, if a user feels the need to relax, the emotion recognition system analyzes the user's emotional state and generates and recommends a playlist of relaxing music based on that analysis. Examples of generation instructions could be text inputs such as, "Generate a music playlist suitable for when the user wants to relax," or "Recommend entertainment videos that match the user's current mood."
[0631] This system allows users to receive personalized images and content adapted to their emotional state in real time, resulting in a more satisfying user experience.
[0632] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0633] Step 1:
[0634] The user accesses the web application and enters generation instructions.
[0635] Input: User text input (e.g., "I want to generate relaxing images")
[0636] Output: Generation instruction data
[0637] Specific action: The user enters text into a designated form on the web application and clicks the "Generate" button.
[0638] Step 2:
[0639] The server begins emotion recognition.
[0640] Input: User's facial expression data, voice data, interaction data such as input speed, etc.
[0641] Output: Emotional state data (e.g., "I want to relax")
[0642] Specific operation: The server uses EmotionEngine to analyze data collected from the camera and microphone to identify the user's emotional state.
[0643] Step 3:
[0644] The server adjusts the generation instructions based on the emotional state.
[0645] Input: Emotional state data, generation instruction data
[0646] Output: Adjusted generation instruction data (e.g., "Relaxing sunset landscape image")
[0647] Specific operation: Based on emotional state data, personalized elements are added to the generation instruction data.
[0648] Step 4:
[0649] The server sends generation instructions to the artificial intelligence model, which then generates the image.
[0650] Input: Adjusted generation instruction data
[0651] Output: Generated image data
[0652] Specific operation: The server sends a prompt message to the generation AI model and generates the specified image (e.g., "generate a relaxing sunset landscape").
[0653] Step 5:
[0654] The server sends the generated image data to the user's terminal for display.
[0655] Input: Generated image data
[0656] Output: Image displayed on the user's terminal
[0657] Specific operation: The server encodes the generated image data and sends it to the user's terminal. The user's terminal decodes the image data and displays it on the web application.
[0658] Step 6:
[0659] The user prints the generated image data.
[0660] Input: Generated image data, user's print instructions
[0661] Output: Printed image
[0662] Specific operation: When the user clicks the "Print" button, the server sends the image data to the printer's API and performs high-quality printing.
[0663] Step 7:
[0664] The user sends generated image data to a display device, which then displays it on a large screen.
[0665] Input: Generated image data, user display instructions
[0666] Output: Image displayed on the display device
[0667] Specific operation: When the user clicks the "Display on TV" button, the server sends the image data to the display device's API and displays it on the large screen in high resolution.
[0668] Step 8:
[0669] After the server finishes generating the image, it displays information about the product, which incorporates artificial intelligence technology, on the user's terminal.
[0670] Input: Generated image data, related product information
[0671] Output: Product information displayed on the user terminal
[0672] Specific operation: The server retrieves relevant AI product information from the generated image data and displays it on the user's terminal. This information includes detailed product information and purchase links.
[0673] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0674] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0675] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0676] [Second Embodiment]
[0677] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0678] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0679] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0680] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0681] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0682] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0683] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0684] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0685] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0686] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0687] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0688] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0689] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. In the following description, the program processing at each step of the system will be explained in natural language, with specific examples.
[0690] Basic configuration
[0691] 1. Preparing the user interface
[0692] Server: Starts a web server and hosts the web application that users can access. This allows users to access a form to enter instructions for image generation.
[0693] 2. Receive instructions for image generation.
[0694] User: Access the web application and enter the content of the image you want to generate as text. Submit this instruction by clicking the "Generate" button.
[0695] Terminal: Receives the entered text data and sends it to the server.
[0696] 3. Image generation using AI
[0697] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "sunset landscape" would be an instruction to generate an image depicting an evening sky and horizon.
[0698] Server: The generative artificial intelligence model generates images and sends the generation results back to the server.
[0699] 4. Displaying the generated image
[0700] Server: Sends the received image data back to the user's terminal.
[0701] Terminal: Receives image data and displays it to the user on the web application.
[0702] 5. Prepare and execute the print job.
[0703] User: Review the generated image and click the "Print" button on the web application.
[0704] Terminal: Triggers a print action and sends image data to the server.
[0705] Server: Sends the received image data to the printer's API and instructs it to print.
[0706] Printer: Print the image according to the instructions. The user can verify that the generated image is printed in high quality.
[0707] 6. Preparing and executing the TV display.
[0708] User: Review the generated image and click the "Display on TV" button.
[0709] Terminal: Triggers a display action and sends image data to the server.
[0710] Server: Sends the received image data to the TV's API and instructs it to display the image.
[0711] Television: Displays images in high resolution. Users can see that the generated images are displayed on a large screen.
[0712] 7. Display of information on AI-equipped products
[0713] Server: After image generation, printing, and display operations are complete, the server provides the user with a web application page displaying detailed information about the product, which incorporates artificial intelligence technology.
[0714] Terminal: Users can view product information they are interested in and perform actions to obtain more detailed information.
[0715] 8. Considering purchasing a Pixel product
[0716] User: View the provided product information, obtain additional information as needed, and consider purchasing the product.
[0717] Server: Provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0718] Specific example
[0719] Example 1: Creating and printing landscape photographs
[0720] User: Enters "I want a landscape photo" into the web application and clicks the "Generate" button.
[0721] Server: Sends user instructions to the generation AI and receives the generated landscape photo data.
[0722] Terminal: Displays received landscape photos.
[0723] User: After reviewing the image, click the "Print" button.
[0724] Server: Sends image data to the printer and starts printing.
[0725] Printer: Prints high-quality landscape photos.
[0726] Example 2: Generating a picture of a cat and displaying it on a TV.
[0727] User: Type "I want to generate a picture of a cat" and click the "Generate" button.
[0728] Server: Sends user instructions to the generation AI and receives the generated cat image data.
[0729] Terminal: Displays the received picture of a cat.
[0730] User: After reviewing the image, click the "Display on TV" button.
[0731] Server: Sends image data to the television and starts displaying it.
[0732] Television: Displays a picture of a cat in high resolution.
[0733] Example 3: Information acquisition for AI-powered products
[0734] User: After completing the image generation experience, view the displayed information page for AI-powered products.
[0735] Server: Provides detailed product information and answers questions via chatbot.
[0736] User: Obtain information to consider purchasing and inquire for further details as needed.
[0737] As described above, this system utilizes generative AI to generate images and provides a series of processes for outputting those images in various ways. Furthermore, by providing users with information about AI-powered products through the generation experience, it becomes possible to create a new purchasing experience.
[0738] The following describes the processing flow.
[0739] Step 1:
[0740] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[0741] Step 2:
[0742] The user accesses the web application and enters the content of the image they want to generate as text. Once the input is complete, they click the "Generate" button.
[0743] Step 3:
[0744] The terminal receives user input and sends text data to the server. The transmitted data is in JSON format.
[0745] Step 4:
[0746] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, it passes the text "sunset landscape" to the "generative AI".
[0747] Step 5:
[0748] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user terminal. The image data is usually encoded in a format such as Base64.
[0749] Step 6:
[0750] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[0751] Step 7:
[0752] After the user reviews the generated image, they click the "Print" button on the web application.
[0753] Step 8:
[0754] The terminal triggers a print action and resends the image data to the server.
[0755] Step 9:
[0756] The server receives the image data and sends it to the printer's API to issue a print command.
[0757] Step 10:
[0758] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[0759] Step 11:
[0760] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[0761] Step 12:
[0762] The device triggers a display action, and the image data is sent to the server again.
[0763] Step 13:
[0764] The server sends image data to the TV's API and issues instructions for display.
[0765] Step 14:
[0766] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[0767] Step 15:
[0768] After the server has finished generating, printing, and displaying images, it provides users with a web application containing detailed product information pages that incorporate AI technology.
[0769] Step 16:
[0770] Users can view product information and perform actions to request additional information as needed.
[0771] Step 17:
[0772] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[0773] (Example 1)
[0774] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0775] Traditional image generation systems have been criticized for their complex user experience, as the process of generating, displaying, and printing images is cumbersome. Furthermore, they lack mechanisms for quickly providing product information related to the generated images, missing opportunities to increase user purchasing intent. Additionally, there are insufficient methods for users to easily ask questions about products and receive answers.
[0776] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0777] In this invention, the server includes means for hosting a web application accessible by a user terminal, means for receiving generation instructions, means for using a generative artificial intelligence model that generates images based on the generation instructions, means for returning and displaying the generated image data to the user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for displaying information about a product equipped with artificial intelligence technology after the image generation is complete, and means for the user to ask questions about the product through a chatbot and receive responses. This allows the user to easily and quickly perform a series of processes from image generation to display, printing, and product information acquisition, resulting in a high-quality user experience.
[0778] A "user terminal" is a device that a user directly operates to access web applications and the internet.
[0779] A "web application" is a software application that operates via the internet or an intranet and is accessible to users through a web browser.
[0780] "Generation instructions" refer to the text or commands that the user enters to generate an image.
[0781] A "generative artificial intelligence model" is a system that includes machine learning algorithms for generating content such as images based on specified input data.
[0782] A "server" is a computer system on a network that hosts web applications and processes requests from users.
[0783] "Image data" refers to the digital data format of a generated image, which is data that can be displayed and printed.
[0784] A "printing device" is a hardware device that receives digital image data and prints it onto physical paper.
[0785] A "display device" is a device that receives image data and displays it on a screen in high resolution.
[0786] "Product information" refers to data including specifications, reviews, and purchase information related to products equipped with artificial intelligence technology.
[0787] A "chatbot" is conversational software that automatically responds to questions from users.
[0788] Basic configuration
[0789] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. This system includes the following elements:
[0790] 1. Preparing the user interface
[0791] Server: Start a web server (e.g., Apache, Nginx) and host a web application using Python and a framework like Django or Flask. This will allow users to access a webpage with a form to input instructions for image generation.
[0792] 2. Receive instructions for image generation.
[0793] User: Access the web application and enter the content of the image you want to generate as text. For example, use a prompt like "Sunset Landscape". This prompt is submitted when you click the "Generate" button.
[0794] Terminal: Captures the entered text data and sends it to the server.
[0795] 3. Image generation using AI
[0796] Server: Analyzes the received text data and sends it as a prompt to a generative AI model (e.g., OpenAI's DALL-E or GPT-4). Examples of prompts include "sunset landscape" and "picture of a cat."
[0797] Server: Receives image data generated by the generation AI model based on instructions and verifies its format (e.g., PNG, JPEG).
[0798] 4. Displaying the generated image
[0799] Server: Sends the generated image data back to the user's terminal.
[0800] Terminal: Renders received image data on a web browser and displays it to the user.
[0801] 5. Prepare and execute the print job.
[0802] User: Review the displayed image and click the "Print" button on the web application.
[0803] Terminal: Triggers a print action and sends image data to the server.
[0804] Server: Sends image data to the printer's API (e.g., HP's Printer API) and issues a print command.
[0805] Printer: Prints high-quality images using the received image data.
[0806] 6. Preparing and executing the TV display.
[0807] User: Check the displayed image and click the "Display on TV" button.
[0808] Terminal: Triggers a display action and sends image data to the server.
[0809] Server: Sends image data to the TV's API (e.g., Chromecast API) and instructs it to display the image.
[0810] Television: Displays images in high resolution.
[0811] 7. Display of information on AI-equipped products
[0812] Server: After the user completes operations such as image generation, printing, and display, the server provides a page in the web application that displays information about the product, which incorporates artificial intelligence technology. The information page includes detailed product specifications, user reviews, and purchase information.
[0813] Terminal: Displays product information on a webpage, allowing users to view details.
[0814] 8. Considering purchasing a Pixel product
[0815] User: Browse the displayed AI product information page and review the details. Enter questions into the chatbot as needed.
[0816] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[0817] Specific examples of operation
[0818] Example 1: Creating and printing landscape photographs
[0819] User: Enters "I want a landscape photo" into the web application and clicks the generate button.
[0820] Server: Sends instructions as prompts to the AI model generating the data (e.g., DALL-E) and receives the generated landscape photos.
[0821] Terminal: The landscape photo is displayed in a web browser, and the user clicks the print button after confirming it.
[0822] Server: Sends image data to the printer's API and issues a print command.
[0823] Printer: Prints landscape photos in high quality.
[0824] Example 2: Generating a picture of a cat and displaying it on a TV.
[0825] User: Enters "I want to generate a picture of a cat" into the web application and clicks the generate button.
[0826] Server: Sends instructions as prompts to the AI model that generates the data, and receives the generated picture of a cat.
[0827] Device: A picture of a cat is displayed on the webpage, and the user clicks the "Show on TV" button after confirming it.
[0828] Server: Sends image data to the TV's API and issues display instructions.
[0829] Television: Displays a picture of a cat in high resolution.
[0830] Example 3: Information acquisition for AI-powered products
[0831] User: After completing the image generation experience, view the displayed AI product information page to learn more details.
[0832] Server: Provides detailed product specifications, user reviews, and purchase information, and answers questions via chatbot functionality.
[0833] User: Obtain information to help with purchase considerations and resolve questions via chatbots or web pages as needed.
[0834] This invention allows users to easily and quickly perform a series of processes from image generation to display, printing, and acquisition of product information. The system aims to provide a high-quality user experience and enhance user purchasing intent.
[0835] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0836] Program processing steps
[0837] Step 1:
[0838] User: Accesses the web application and enters instructions for image generation. Specifically, the user enters prompt text such as "Sunset Landscape" or "Picture of a Cat" into the text input field and clicks the "Generate" button.
[0839] Input: Prompt text (e.g., "Sunset scenery")
[0840] Output: Text data sent by user interaction
[0841] Step 2:
[0842] Terminal: Sends the entered text data to the server. Specifically, when the user clicks the "Generate" button, that text data is sent to the server as an HTTP request.
[0843] Input: Text data entered by the user
[0844] Output: Text data sent to the server
[0845] Step 3:
[0846] Server: Analyzes the received text data and sends it as a prompt to the generative AI model. Natural language processing (NLP) techniques are used for analysis, and API requests are generated to the generative AI model (e.g., OpenAI's DALL-E or GPT-4).
[0847] Input: Text data received from the device
[0848] Output: API request to the generated AI model
[0849] Step 4:
[0850] Generative AI model: Generates images based on text data. The generated image data is returned to the server.
[0851] Input: Prompt message sent from the server (API request)
[0852] Output: Generated image data
[0853] Step 5:
[0854] Server: Retrieves image data received from the generated AI model and sends it back to the user's terminal. Here, it checks the image data format (e.g., PNG, JPEG) and encodes it in the appropriate format.
[0855] Input: Image data returned from a generative AI model
[0856] Output: Image data to be sent to the user's terminal.
[0857] Step 6:
[0858] Terminal: Renders received image data on a web browser and displays it to the user.
[0859] Input: Image data returned from the server
[0860] Output: Image displayed in the web browser
[0861] Step 7:
[0862] User: Review the generated image and click the "Print" or "Display on TV" button on the web application.
[0863] Input: Image data displayed in a web browser
[0864] Output: User instructions for printing or displaying.
[0865] Step 8:
[0866] Terminal: Triggers user actions (print or view) and sends image data to the server.
[0867] Input: User's print or display instructions
[0868] Output: Image data sent to the server
[0869] Step 9:
[0870] Server: Sends image data to the corresponding device's API to instruct it to print or display. For printing, it sends the data to the printer's API (e.g., HP's Printer API), and for display, it sends the data to the TV's API (e.g., Chromecast API).
[0871] Input: Image data sent from the device
[0872] Output: Image data sent to a printer or television.
[0873] Step 10:
[0874] Printer or television: Receives image data and displays or prints it on the respective device. A printer prints the image on paper, while a television displays the image in high resolution.
[0875] Input: Image data sent from the server
[0876] Output: Printed image or displayed image
[0877] Step 11:
[0878] Server: After the user completes the image generation, printing, and display operations, the server provides a page in the web application that displays product information powered by artificial intelligence technology.
[0879] Input: User operation completion information
[0880] Output: Product information display page
[0881] Step 12:
[0882] User: View information about AI-powered products on a webpage and check the details. Enter questions into the chatbot as needed.
[0883] Input: Viewing product information page
[0884] Output: Question sent to the chatbot
[0885] Step 13:
[0886] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[0887] Input: Question from a user
[0888] Output: Chatbot or human response
[0889] (Application Example 1)
[0890] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0891] Conventional image generation systems had a cumbersome process for users to apply their generated images to specific products and lacked intuitive preview functionality. As a result, users could not visually confirm how the generated images would appear on the actual product, making it difficult to achieve a satisfactory purchasing experience. Furthermore, the means of transmitting the generated images to various output devices for appropriate display and printing were limited.
[0892] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0893] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for applying the generated images to a specific product, means for displaying the result of applying the generated images to a specific product as a three-dimensional preview, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This allows the user to intuitively apply the generated images to a specific product and confirm them in a three-dimensional preview. Furthermore, the generated images can be easily displayed and printed on various output devices, improving the user's purchasing experience.
[0894] A "means for receiving generation instructions" refers to a device or system that receives input from a user to instruct image generation.
[0895] "Means using a generative artificial intelligence model that generates images based on generation instructions" refers to devices or systems that utilize artificial intelligence technology to generate images based on user instructions.
[0896] "Means for displaying generated images on a user's terminal" refers to devices or systems for displaying generated image data on a user's terminal.
[0897] "Means for transmitting generated image data to a printing device and performing printing" refers to a device or system for transmitting generated image data to a printing device and printing the image using the printing device.
[0898] "Means for transmitting generated image data to a display device and displaying it" refers to a device or system for transmitting generated image data to a display device and displaying the image on the display device.
[0899] "Means for applying a generated image to a specific product" refers to a device or system for applying a generated image to a desired product.
[0900] "Means for displaying the result of applying a generated image to a specific product as a three-dimensional preview" refers to a device or system for applying a generated image to a specific product and visualizing and displaying the result in three dimensions.
[0901] "Means for displaying information about a product equipped with artificial intelligence technology after the generation of the aforementioned image is completed" refers to a device or system for providing information about a product equipped with artificial intelligence technology after the generation process is completed.
[0902] The following describes the specific system configuration and program processing for implementing this invention. In particular, it describes the detailed process of generating images using a generative AI model and applying them to a product.
[0903] Basic System Configuration
[0904] 1. Preparing the user interface
[0905] Server: Starts a web server and hosts a web application for users to access. This web application includes a form for entering instructions for image generation.
[0906] 2. Receive instructions for image generation.
[0907] User: Access the web application and enter the content of the image you want to generate as text. Submit this text instruction by clicking the "Generate" button.
[0908] 3. Image generation using AI
[0909] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "Colorful geometric patterns" becomes an instruction to generate an image based on its content.
[0910] Server: The generative artificial intelligence model generates images and sends the results back to the server.
[0911] 4. Displaying the generated image
[0912] Server: Sends the received image data back to the user's terminal.
[0913] Terminal: Receives image data and displays it to the user on the web application.
[0914] 5. Application of images to products
[0915] User: Review the generated image and select the option to apply it to a specific product (e.g., T-shirt, cup, etc.).
[0916] Server: Uses image data to place and apply images to specified products.
[0917] Terminal: Displays a 3D preview of the generated product for the user to review.
[0918] 6. Prepare and execute printing.
[0919] User: After reviewing the generated image, decide to order the product and click the "Print" button.
[0920] Server: Receives print instructions and sends the generated image data to the printing device.
[0921] Printing device: Follow the instructions and print the generated image onto the product.
[0922] 7. Preparing and executing the TV display.
[0923] User: After reviewing the generated image, click the "Display on TV" button.
[0924] Server: Triggers a display action and sends image data to the display device.
[0925] Display device: Displays images in high resolution.
[0926] 8. Display of information on AI-equipped products
[0927] Server: After image generation, printing, and display operations are complete, the server displays a page on the web application that provides users with detailed information about the product, which incorporates artificial intelligence technology.
[0928] Hardware and software to be used
[0929] hardware
[0930] User's smartphone
[0931] Server (e.g., cloud service)
[0932] printing device
[0933] Display device (home television)
[0934] software
[0935] Frontend: React
[0936] Backend: Flask (Python)
[0937] Image generation: OpenAI API
[0938] Specific example
[0939] Example of a prompt
[0940] "Colorful geometric patterns"
[0941] "Spacescape"
[0942] "Cute cat illustration"
[0943] For example, a user accesses a web application, enters "colorful geometric patterns," and clicks the "Generate" button. The server, upon receiving this instruction, uses a generative AI model to generate an image based on the specified content. The generated image is first displayed on the user's terminal. Next, the user applies the image to a specific product and checks it in a 3D preview. If the user is satisfied with this preview, they can send the image to a printing device to print it on the product. The generated image can also be displayed on a home television. Finally, after the image generation experience, the user is provided with detailed information about products equipped with artificial intelligence technology.
[0944] In summary, this invention provides a system that allows users to intuitively generate images and quickly apply them to products for visualization. Furthermore, user satisfaction can be enhanced by displaying or printing the generated images on various output devices.
[0945] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0946] Step 1:
[0947] The user enters a prompt.
[0948] The user accesses the web application and enters a text prompt for image generation (e.g., "Colorful geometric pattern"). Once the input is complete, they click the "Generate" button.
[0949] Input: Text prompt (e.g., "Colorful geometric patterns")
[0950] Output: Generation instructions
[0951] Specific operation: The user enters a prompt message in the text input field and clicks a button, which sends the instruction to the server.
[0952] Step 2:
[0953] The server receives the generation instruction.
[0954] The server receives the prompt message sent by the user and parses it.
[0955] Input: Generation Instructions
[0956] Output: Parsed prompt message
[0957] Specific operation: The server receives an HTTP request and extracts the prompt text.
[0958] Step 3:
[0959] The server sends a request to the generated AI model.
[0960] Based on the parsed prompt text, the server sends an image generation request to the AI model.
[0961] Input: Parsed prompt message
[0962] Output: Request for generated AI model
[0963] Specific operation: The server generates a prompt message and sends it to the AI model's API in the appropriate format.
[0964] Step 4:
[0965] The generative AI model generates images.
[0966] The generative AI model generates an image based on the prompt text and sends the result back to the server.
[0967] Input: Request for generated AI model
[0968] Output: Generated image data
[0969] Specific operation: The generative AI model performs internal processing, generates an image based on the prompt text, and sends that data back to the server.
[0970] Step 5:
[0971] The server receives the generated image and sends it to the user's terminal.
[0972] The server receives the generated image data and sends it back to the user's terminal.
[0973] Input: Generated image data
[0974] Output: Sending image data to the user terminal
[0975] Specific operation: The server sends the received image data to the user's terminal as an HTTP response.
[0976] Step 6:
[0977] Display images on the user's terminal.
[0978] The user terminal displays the received image data on the web application.
[0979] Input: Image data sent to the user terminal
[0980] Output: Image display on a web application
[0981] Specific operation: The user's device browser renders the image data and displays it on the screen.
[0982] Step 7:
[0983] Users apply images to the product.
[0984] The user reviews the generated image and selects the option to apply it to a specific product (e.g., a T-shirt, a cup).
[0985] Input: Generated image and product selection
[0986] Output: Image applied to the product
[0987] Specific operation: The user clicks a product selection option and sends an apply request to the server.
[0988] Step 8:
[0989] The server applies the image to the product.
[0990] The server applies the generated image to a specific product and generates the corresponding data.
[0991] Input: Generated image and product selection information
[0992] Output: Image data applied to the product
[0993] Specific operation: The server places images into the specified product template and generates the applied data.
[0994] Step 9:
[0995] The server sends a 3D preview to the user terminal.
[0996] The server sends the generated product data to the user terminal in a three-dimensional preview format.
[0997] Input: Image data applied to the product
[0998] Output: Sending a 3D preview to the user's terminal
[0999] Specific operation: The server generates data for the 3D preview and sends it to the user's terminal as an HTTP response.
[1000] Step 10:
[1001] The user terminal displays a 3D preview.
[1002] The user terminal displays the received 3D preview on the web application.
[1003] Input: Data for 3D preview
[1004] Output: Three-dimensional preview display on screen
[1005] Specific operation: The user's browser renders a 3D preview and displays it on the screen.
[1006] Step 11:
[1007] The user issues a print command.
[1008] The user reviews the 3D preview, and if satisfied, clicks the "Print" button to initiate printing.
[1009] Input: User's print instructions
[1010] Output: Print request to server
[1011] Specific action: The user clicks a button and sends a print command to the server.
[1012] Step 12:
[1013] The server sends image data to the printer.
[1014] The server receives the print command and sends the generated image data to the printing device.
[1015] Input: User's print instructions and image data.
[1016] Output: Print instructions to the printer
[1017] Specific operation: The server sends image data to the printer's API and issues a print command.
[1018] Step 13:
[1019] The printing device prints the image.
[1020] The printing device prints an image onto the specified product based on the received image data.
[1021] Input: Print instructions to the printer and image data.
[1022] Output: Printed product
[1023] Specific operation: The printing device prints an image onto the product based on the data it receives.
[1024] Step 14:
[1025] The user displays an image on the display device.
[1026] The user clicks the "Display on TV" button to display the generated image on a display device such as a home television.
[1027] Input: User's display instructions
[1028] Output: Display request to the server
[1029] Specific operation: The user clicks a button and sends a display instruction to the server.
[1030] Step 15:
[1031] The server sends image data to the display device.
[1032] The server receives the display instruction and sends the generated image data to the display device.
[1033] Input: User display instructions and image data
[1034] Output: Display instructions for the display device.
[1035] Specific operation: The server sends image data to the display device's API and issues a display instruction.
[1036] Step 16:
[1037] The display device displays an image.
[1038] The display device displays the image in high resolution based on the received image data.
[1039] Input: Display instructions and image data for the display device.
[1040] Output: Displayed image
[1041] Specific operation: The display device displays an image on the screen based on the data it receives.
[1042] Step 17:
[1043] The server displays information about products equipped with AI technology.
[1044] After the server completes image generation, printing, and display operations, it displays a page on the web application that provides users with detailed information about the AI-powered product.
[1045] Input: User's operation complete
[1046] Output: Information on products equipped with AI technology
[1047] Specific action: The server updates the content of the webpage and provides the user with detailed information.
[1048] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1049] This invention relates to a generative AI system that combines an emotion engine to recognize user emotions. This system receives image generation instructions from the user, generates images using a generative artificial intelligence model, and outputs the generated images to a printer or display device. Furthermore, it displays information about products equipped with artificial intelligence technology after image generation. The system incorporates an emotion engine that recognizes the user's emotional state and adjusts generation instructions based on that state.
[1050] Basic configuration
[1051] 1. Preparing the user interface
[1052] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[1053] 2. Initiation of emotion recognition
[1054] The user accesses the web application and clicks a specific button to initiate emotion recognition.
[1055] The device collects user interaction and input data and sends it to the emotion engine.
[1056] 3. Analysis of emotional state
[1057] The server uses an emotion engine to analyze the user's emotional state. For example, it identifies emotions based on the user's input speed, input content, and facial expression data (if acquired using a camera).
[1058] 4. Receive instructions for image generation.
[1059] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[1060] The emotion engine takes the user's emotional state into consideration and makes adaptive adjustments to the generation instructions.
[1061] The terminal sends the adjusted generation instruction data to the server.
[1062] 5. Image generation using AI
[1063] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, based on the input text "Sunset Landscape," it generates an image that contains emotionally positive elements.
[1064] 6. Displaying the generated image
[1065] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user's terminal.
[1066] The device receives image data and displays it to the user on the web application. The user then views the generated image on the screen.
[1067] 7. Prepare and execute the print job.
[1068] The user reviews the generated image and clicks the "Print" button on the web application.
[1069] The terminal triggers a print action and resends the image data to the server.
[1070] The server receives the image data and sends it to the printer's API to issue a print command.
[1071] The printer prints the image in high quality according to the instructions. The user receives the printed image.
[1072] 8. Preparing and executing the TV display.
[1073] After the user reviews the generated image, they click the "Display on TV" button.
[1074] The device triggers a display action, and the image data is sent to the server again.
[1075] The server sends image data to the TV's API and issues instructions for display.
[1076] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[1077] 9. Information display for AI-equipped products
[1078] After the server has finished generating, printing, and displaying images, it provides users with detailed information about the AI-powered product via a web application.
[1079] The device displays the product information page to the user.
[1080] 10. Considering purchasing a Pixel product
[1081] Users view product information and, if interested, take actions to request additional information.
[1082] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1083] Specific example
[1084] Example 1: Emotion-driven landscape photography and printing
[1085] The user enters "I want a landscape photo" and clicks the "Generate" button.
[1086] The emotion engine analyzes user input and interaction data to identify positive emotional states.
[1087] The server sends an instruction to the AI model to generate a "positive sunset landscape."
[1088] The server receives the generated landscape photo data and sends it to the terminal.
[1089] The device displays landscape photos it has received.
[1090] After the user reviews the image, they click the "Print" button.
[1091] The server sends the image data to the printer and starts printing.
[1092] The printer prints high-quality landscape photographs.
[1093] Example 2: Generation and display of cat images based on emotions
[1094] The user enters "I want to generate a picture of a cat" and clicks the "Generate" button.
[1095] The emotion engine analyzes the user's interaction data and determines that the user is slightly tired.
[1096] The server sends an instruction to the AI model to generate a "relaxing picture of a cat."
[1097] The server receives the generated cat image data and sends it to the terminal.
[1098] The device displays the image of a cat it received.
[1099] After the user reviews the image, they click the "Display on TV" button.
[1100] The server sends the image data to the television and begins displaying it.
[1101] The TV displays a picture of a cat in high resolution.
[1102] Example 3: Information retrieval for AI-powered products based on emotions
[1103] After the user completes the image generation experience, they can view information pages about AI-powered products on the web application.
[1104] The emotion engine recognizes the user's level of excitement and prioritizes displaying highly relevant product information.
[1105] The server provides detailed product information, and a chatbot answers questions.
[1106] Users can obtain information to consider purchasing and inquire about details as needed.
[1107] In this way, a system is realized that takes user emotions into consideration and provides a consistent service from image generation using generative AI to output methods and product information acquisition. By incorporating an emotion engine, the user experience can be further personalized and satisfaction can be increased.
[1108] The following describes the processing flow.
[1109] Step 1:
[1110] The server starts up a web server and hosts a web application for users to access. This allows users to access a form to enter instructions for image generation.
[1111] Step 2:
[1112] The user accesses a web application and clicks a specific button to initiate emotion recognition. This action causes the emotion engine to begin collecting data about the user.
[1113] Step 3:
[1114] The device collects user interaction data (such as input speed, input content, and in some cases, facial recognition data from the webcam) and sends it to the emotion engine.
[1115] Step 4:
[1116] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the collected data to determine whether the user is in a positive, negative, relaxed, or other emotional state.
[1117] Step 5:
[1118] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[1119] Step 6:
[1120] The emotion engine considers the user's emotional state and makes adaptive adjustments to the "generate" instructions. For example, if the user is in a positive state, it adds instructions to include vibrant colors and bright scenes in the generated image.
[1121] Step 7:
[1122] The terminal sends the adjusted generation instruction data to the server. The transmitted data includes text as well as adjustment information for the emotion engine.
[1123] Step 8:
[1124] The server analyzes the received text data and sentiment adjustment data, and sends an image generation request to the generative artificial intelligence model. For example, when providing the generative AI model with the text "sunset landscape," instructions are also given to include positive elements.
[1125] Step 9:
[1126] The server receives the generated image data returned from the generated artificial intelligence model and sends it to the user's terminal. The image data is usually encoded in a format such as Base64.
[1127] Step 10:
[1128] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[1129] Step 11:
[1130] After the user reviews the generated image, they click the "Print" button on the web application.
[1131] Step 12:
[1132] The terminal triggers a print action and resends the image data to the server.
[1133] Step 13:
[1134] The server receives the image data and sends it to the printer's API to issue a print command.
[1135] Step 14:
[1136] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[1137] Step 15:
[1138] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[1139] Step 16:
[1140] The device triggers a display action, and the image data is sent to the server again.
[1141] Step 17:
[1142] The server sends image data to the TV's API and issues instructions for display.
[1143] Step 18:
[1144] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[1145] Step 19:
[1146] After the server completes the image generation, printing, and display operations, it provides users with detailed information about the product, which incorporates artificial intelligence technology, through a web application.
[1147] Step 20:
[1148] Users view the displayed product information and, if interested, take action to request additional information.
[1149] Step 21:
[1150] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1151] (Example 2)
[1152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1153] Conventional image generation systems often fail to consider the user's emotional state when issuing generation instructions, resulting in generated images that do not always match the user's expectations or feelings. Furthermore, the limited output formats and methods for generated images posed a challenge, hindering user convenience.
[1154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1155] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for adjusting the generation instructions using an emotion engine that recognizes the user's emotional state, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This makes it possible to generate images that take the user's emotional state into consideration, thereby increasing user satisfaction. In addition, the generated images can be output to a variety of devices, improving user convenience.
[1156] "Generation instructions" are data entered by the user to specify the content and conditions of the image that should be generated.
[1157] A "generative artificial intelligence model" is an artificial intelligence algorithm or system that generates images based on instructions or data input by a user.
[1158] A "user terminal" refers to a device such as a computer or smartphone that a user operates, and which provides an interface for image generation through a web application.
[1159] A "printing device" is a hardware device used to physically print generated images, and primarily refers to a printer.
[1160] A "display device" is a hardware device used to visually present generated images to a user, and includes televisions and monitors.
[1161] An "emotion engine" is software or a system that analyzes user interaction data and input to identify emotional states and adjusts generation instructions based on those results.
[1162] "Information about products equipped with artificial intelligence technology" refers to detailed information about products incorporating artificial intelligence technology, which is provided to the user after the display of the generated image is complete.
[1163] This invention relates to an image generation system that incorporates an emotion engine to recognize user emotions. The system receives image generation instructions from the user and generates images using a generation artificial intelligence model. The generated images are then output to a printer or display device, and information about the product, which incorporates artificial intelligence technology, is displayed after image generation.
[1164] composition
[1165] The system of the present invention has the following components.
[1166] 1. Preparing the user interface
[1167] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is implemented using HTML / CSS / JavaScript.
[1168] 2. Initiation of emotion recognition
[1169] The user clicks the "Start Emotion Recognition" button in the web application. The device captures the click event and prepares to send interaction data to the emotion engine. This is done using services such as Microsoft Azure's Emotion API.
[1170] 3. Analysis of emotional state
[1171] The server sends the received interaction data to the emotion engine, which then analyzes the user's emotional state. For example, it analyzes input content, input speed, and facial expression data captured by the camera.
[1172] 4. Receive instructions for image generation.
[1173] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the emotional state and makes adaptive adjustments to the generation instructions. The device then sends the adjusted generation instruction data to the server.
[1174] 5. Image generation using AI
[1175] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[1176] 6. Displaying the generated image
[1177] The server receives the generated image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[1178] 7. Prepare and execute the print job.
[1179] The user reviews the generated image and clicks the "Print" button. The device resends the image data to the server. The server sends the data to the printer's API (such as Google Cloud Print) and prints the image.
[1180] 8. Preparing and executing the TV display.
[1181] After the user reviews the generated image, they click the "Display on TV" button. The device sends the image data to the server, which then sends it to the TV's API (such as Google Cast or Samsung Smart View API) for display.
[1182] 9. Information display for AI-equipped products
[1183] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[1184] 10. Considering purchasing a Pixel product
[1185] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[1186] Specific example
[1187] Example 1: Creating and printing landscape photographs based on emotions.
[1188] User: Enter "I want a landscape photo" and click the "Generate" button.
[1189] Emotion Engine: Analyzes user input and emotional states to identify positive emotional states.
[1190] Server: Sends the instruction "positive sunset scenery" to the generating artificial intelligence model.
[1191] Server: Receives the generated image data and sends it to the terminal.
[1192] Terminal: Receives and displays image data.
[1193] User: Click the "Print" button.
[1194] Server: Sends image data to the printer API and starts printing.
[1195] Printer: Prints high-quality landscape photos.
[1196] Example 2: Generation of cat images based on emotions and display on television
[1197] User: Enter "I want to generate a picture of a cat" and click the "Generate" button.
[1198] Emotion Engine: Analyzes user interaction data to recognize fatigue levels.
[1199] Server: Instructs the artificial intelligence model to generate a "relaxing picture of a cat".
[1200] Server: Receives the generated cat image data and sends it to the terminal.
[1201] Terminal: Receives and displays image data.
[1202] User: Click the "Display on TV" button.
[1203] Server: Sends image data to the TV API and starts displaying it on the TV.
[1204] Television: Displays a picture of a cat in high resolution.
[1205] Example 3: Information acquisition for AI-powered products based on emotions
[1206] User: After completing the image generation experience, they browse the information page for AI-powered products in the web application.
[1207] Emotion Engine: Recognizes the user's emotional state and prioritizes displaying highly relevant product information.
[1208] Server: Provides detailed product information and answers questions via chatbot.
[1209] User: Obtain information to consider purchasing and inquire for further details as needed.
[1210] In this way, by incorporating an emotion engine, we can personalize the user experience and create a system that enhances satisfaction. The collaboration between the generative AI model and the emotion engine enables the generation and output of appropriate images that take into account the user's emotional state.
[1211] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1212] Step 1: Preparing the User Interface
[1213] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is constructed using HTML, CSS, and JavaScript. Users access it through a web browser and are presented with a form to input instructions for image generation.
[1214] Input: User access request
[1215] Output: Display of a form for entering image generation instructions.
[1216] Specific operation: The server starts up, and the web application waits for user requests. When a user accesses the web application, an instruction input form is displayed.
[1217] Step 2: Initiating Emotion Recognition
[1218] The user clicks the "Start emotion recognition" button in the web application. The device captures the click event and collects user interaction data (facial expressions, typing speed, text input, etc.).
[1219] Input: User click event
[1220] Output: Start of interaction data collection
[1221] Specific actions: The user clicks a button, the device detects the click event and turns on the camera. The device prepares to collect facial expression data and text input data.
[1222] Step 3: Analyze your emotional state
[1223] The server sends interaction data from the terminal to an emotion engine (for example, Microsoft Azure's Emotion API) to analyze the user's emotional state.
[1224] Input: Interaction data from the device
[1225] Output: Sentiment analysis results from the emotion engine
[1226] Specific operation: The server receives data sent from the terminal and sends it to the emotion engine. The emotion engine identifies the emotional state and returns the result to the server.
[1227] Step 4: Receive instructions for image generation.
[1228] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the user's emotional state and adaptively adjusts the generation instructions. The device then sends the adjusted generation instruction data to the server.
[1229] Input: User's image generation instructions
[1230] Output: Adjusted generation instruction data
[1231] Specific operation: The user enters the image content as text and clicks the "Generate" button. The device sends this input data to the server, which then adjusts the instructions via the emotion engine.
[1232] Step 5: Image generation using AI
[1233] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[1234] Input: Adjusted generation instruction data
[1235] Output: Generated image data
[1236] Specific operation: The server analyzes the text data and sends it to the generative AI model as a prompt. The generative AI model generates an image and sends it back to the server.
[1237] Step 6: Displaying the generated image
[1238] The server receives the generated image data returned from the generated AI model and sends it to the terminal. The terminal receives the image data and displays it to the user on the web application.
[1239] Input: Generated image data
[1240] Output: Image display on a web application
[1241] Specific operation: The server receives image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[1242] Step 7: Prepare and execute print
[1243] The user reviews the generated image and clicks the "Print" button. The device triggers the print action and resends the image data to the server. The server sends the image data to the printer API and starts printing.
[1244] Input: User's print instructions
[1245] Output: Printed image
[1246] Specific operation: The user reviews the image and clicks the "Print" button. The device sends the image data to the server, and the server sends a print command to the printer API. The printer prints the image in high quality.
[1247] Step 8: Prepare and run the TV display.
[1248] After the user reviews the generated image, they click the "Display on TV" button. The device triggers the display action and sends the image data to the server again. The server sends the image data to the TV's API and starts displaying it.
[1249] Input: User's display instructions
[1250] Output: Image displayed on the TV
[1251] Specific steps: The user views the image and clicks the "Display on TV" button. The device sends the image data to the server, and the server sends a display command to the TV API. The TV displays the image in high resolution.
[1252] Step 9: Displaying information about AI-powered products
[1253] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[1254] Input: State after operation is complete
[1255] Output: Product Information Page
[1256] Specific operation: The server adds product information to the web application, and the terminal displays the product information page to the user.
[1257] Step 10: Consider purchasing a Pixel product
[1258] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[1259] Input: User's question
[1260] Output: Automated chatbot response and transfer to a human resource representative.
[1261] Specific operation: The user views product information and enters a question into the chatbot. The server provides an automated response and, if necessary, transfers the user to a human representative.
[1262] (Application Example 2)
[1263] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1264] Conventional image generation systems do not provide personalized image generation or content recommendations based on the user's emotional state. This results in a limited user experience and potentially lower satisfaction. Furthermore, there is a need to integrate emotion recognition technology into image generation and content recommendations to provide services that are more emotionally resonant.
[1265] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1266] In this invention, the server includes means for receiving generation instructions, means for using a generative artificial intelligence model to generate images based on the generation instructions, means for analyzing the user's emotional state using emotion recognition means, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, and means for displaying information about products equipped with artificial intelligence technology after the image generation is complete. This enables personalized image generation and content recommendations according to the user's emotional state.
[1267] A "generation instruction" refers to a request regarding images or content that the user will generate.
[1268] A "generative artificial intelligence model" is a part of an artificial intelligence system that includes algorithms and databases for generating images and content based on user instructions and data.
[1269] "Emotion recognition means" refers to a part of a system that includes hardware and software for analyzing the user's emotional state from facial expressions, tone of voice, input speed, etc.
[1270] "Generated image data" refers to the digital representation of image information generated by a generative artificial intelligence model.
[1271] A "user terminal" is a device used to display or print generated images and recommended content.
[1272] A "printing device" is a machine used to print generated image data into a physical format.
[1273] A "display device" is a screen or monitor used to visually display generated image data or recommended content.
[1274] "Products equipped with artificial intelligence technology" refer to goods and services that provide functions using artificial intelligence technologies such as image generation and emotion recognition.
[1275] "Personalized content" refers to content that is customized for a specific user based on their emotional state and preferences.
[1276] The system for implementing this invention is an emotion recognition and content generation system that recognizes the user's emotions and generates and recommends personalized content accordingly.
[1277] This system includes a means of receiving generation instructions from users. A server hosts a web application with a user interface, from which users input instructions for generating images and content. These instructions include the content of the images and content desired by the user.
[1278] The server uses a generative artificial intelligence model to generate images and content based on generation instructions. This model has the ability to generate high-quality images and content based on input such as text prompts.
[1279] Next, the emotion recognition system analyzes the user's emotional state. This system includes a camera and microphone to capture the user's facial expressions, voice tone, input speed, etc., and software to analyze the input data. EmotionEngine, using TensorFlow, identifies the emotional state.
[1280] The server displays the generated image data and content on the user's terminal. The user's terminal is a compatible device such as a smartphone, smart glasses, or head-mounted display.
[1281] Furthermore, it includes means for sending the generated image data to a printing device and printing it in physical form. To achieve this, the server uses the printer's API to send the image data and issue print commands.
[1282] In addition, the system includes means for transmitting the generated image data to a display device and displaying it on a large screen or high-resolution display.
[1283] Finally, the server displays information about the AI-powered product after image generation is complete. This allows users to obtain detailed information and purchase options for the AI-powered product associated with the generated content.
[1284] For example, if a user feels the need to relax, the emotion recognition system analyzes the user's emotional state and generates and recommends a playlist of relaxing music based on that analysis. Examples of generation instructions could be text inputs such as, "Generate a music playlist suitable for when the user wants to relax," or "Recommend entertainment videos that match the user's current mood."
[1285] This system allows users to receive personalized images and content adapted to their emotional state in real time, resulting in a more satisfying user experience.
[1286] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1287] Step 1:
[1288] The user accesses the web application and enters generation instructions.
[1289] Input: User text input (e.g., "I want to generate relaxing images")
[1290] Output: Generation instruction data
[1291] Specific action: The user enters text into a designated form on the web application and clicks the "Generate" button.
[1292] Step 2:
[1293] The server begins emotion recognition.
[1294] Input: User's facial expression data, voice data, interaction data such as input speed, etc.
[1295] Output: Emotional state data (e.g., "I want to relax")
[1296] Specific operation: The server uses EmotionEngine to analyze data collected from the camera and microphone to identify the user's emotional state.
[1297] Step 3:
[1298] The server adjusts the generation instructions based on the emotional state.
[1299] Input: Emotional state data, generation instruction data
[1300] Output: Adjusted generation instruction data (e.g., "Relaxing sunset landscape image")
[1301] Specific operation: Based on emotional state data, personalized elements are added to the generation instruction data.
[1302] Step 4:
[1303] The server sends generation instructions to the artificial intelligence model, which then generates the image.
[1304] Input: Adjusted generation instruction data
[1305] Output: Generated image data
[1306] Specific operation: The server sends a prompt message to the generation AI model and generates the specified image (e.g., "generate a relaxing sunset landscape").
[1307] Step 5:
[1308] The server sends the generated image data to the user's terminal for display.
[1309] Input: Generated image data
[1310] Output: Image displayed on the user's terminal
[1311] Specific operation: The server encodes the generated image data and sends it to the user's terminal. The user's terminal decodes the image data and displays it on the web application.
[1312] Step 6:
[1313] The user prints the generated image data.
[1314] Input: Generated image data, user's print instructions
[1315] Output: Printed image
[1316] Specific operation: When the user clicks the "Print" button, the server sends the image data to the printer's API and performs high-quality printing.
[1317] Step 7:
[1318] The user sends generated image data to a display device, which then displays it on a large screen.
[1319] Input: Generated image data, user display instructions
[1320] Output: Image displayed on the display device
[1321] Specific operation: When the user clicks the "Display on TV" button, the server sends the image data to the display device's API and displays it on the large screen in high resolution.
[1322] Step 8:
[1323] After the server finishes generating the image, it displays information about the product, which incorporates artificial intelligence technology, on the user's terminal.
[1324] Input: Generated image data, related product information
[1325] Output: Product information displayed on the user terminal
[1326] Specific operation: The server retrieves relevant AI product information from the generated image data and displays it on the user's terminal. This information includes detailed product information and purchase links.
[1327] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1328] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1329] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1330] [Third Embodiment]
[1331] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1332] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1333] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1334] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1335] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1336] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1337] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1338] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1339] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1340] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1341] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1342] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1343] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. In the following description, the program processing at each step of the system will be explained in natural language, with specific examples.
[1344] Basic configuration
[1345] 1. Preparing the user interface
[1346] Server: Starts a web server and hosts the web application that users can access. This allows users to access a form to enter instructions for image generation.
[1347] 2. Receive instructions for image generation.
[1348] User: Access the web application and enter the content of the image you want to generate as text. Submit this instruction by clicking the "Generate" button.
[1349] Terminal: Receives the entered text data and sends it to the server.
[1350] 3. Image generation using AI
[1351] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "sunset landscape" would be an instruction to generate an image depicting an evening sky and horizon.
[1352] Server: The generative artificial intelligence model generates images and sends the generation results back to the server.
[1353] 4. Displaying the generated image
[1354] Server: Sends the received image data back to the user's terminal.
[1355] Terminal: Receives image data and displays it to the user on the web application.
[1356] 5. Prepare and execute the print job.
[1357] User: Review the generated image and click the "Print" button on the web application.
[1358] Terminal: Triggers a print action and sends image data to the server.
[1359] Server: Sends the received image data to the printer's API and instructs it to print.
[1360] Printer: Print the image according to the instructions. The user can verify that the generated image is printed in high quality.
[1361] 6. Preparing and executing the TV display.
[1362] User: Review the generated image and click the "Display on TV" button.
[1363] Terminal: Triggers a display action and sends image data to the server.
[1364] Server: Sends the received image data to the TV's API and instructs it to display the image.
[1365] Television: Displays images in high resolution. Users can see that the generated images are displayed on a large screen.
[1366] 7. Display of information on AI-equipped products
[1367] Server: After image generation, printing, and display operations are complete, the server provides the user with a web application page displaying detailed information about the product, which incorporates artificial intelligence technology.
[1368] Terminal: Users can view product information they are interested in and perform actions to obtain more detailed information.
[1369] 8. Considering purchasing a Pixel product
[1370] User: View the provided product information, obtain additional information as needed, and consider purchasing the product.
[1371] Server: Provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1372] Specific example
[1373] Example 1: Creating and printing landscape photographs
[1374] User: Enters "I want a landscape photo" into the web application and clicks the "Generate" button.
[1375] Server: Sends user instructions to the generation AI and receives the generated landscape photo data.
[1376] Terminal: Displays received landscape photos.
[1377] User: After reviewing the image, click the "Print" button.
[1378] Server: Sends image data to the printer and starts printing.
[1379] Printer: Prints high-quality landscape photos.
[1380] Example 2: Generating a picture of a cat and displaying it on a TV.
[1381] User: Type "I want to generate a picture of a cat" and click the "Generate" button.
[1382] Server: Sends user instructions to the generation AI and receives the generated cat image data.
[1383] Terminal: Displays the received picture of a cat.
[1384] User: After reviewing the image, click the "Display on TV" button.
[1385] Server: Sends image data to the television and starts displaying it.
[1386] Television: Displays a picture of a cat in high resolution.
[1387] Example 3: Information acquisition for AI-powered products
[1388] User: After completing the image generation experience, view the displayed information page for AI-powered products.
[1389] Server: Provides detailed product information and answers questions via chatbot.
[1390] User: Obtain information to consider purchasing and inquire for further details as needed.
[1391] As described above, this system utilizes generative AI to generate images and provides a series of processes for outputting those images in various ways. Furthermore, by providing users with information about AI-powered products through the generation experience, it becomes possible to create a new purchasing experience.
[1392] The following describes the processing flow.
[1393] Step 1:
[1394] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[1395] Step 2:
[1396] The user accesses the web application and enters the content of the image they want to generate as text. Once the input is complete, they click the "Generate" button.
[1397] Step 3:
[1398] The terminal receives user input and sends text data to the server. The transmitted data is in JSON format.
[1399] Step 4:
[1400] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, it passes the text "sunset landscape" to the "generative AI".
[1401] Step 5:
[1402] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user terminal. The image data is usually encoded in a format such as Base64.
[1403] Step 6:
[1404] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[1405] Step 7:
[1406] After the user reviews the generated image, they click the "Print" button on the web application.
[1407] Step 8:
[1408] The terminal triggers a print action and resends the image data to the server.
[1409] Step 9:
[1410] The server receives the image data and sends it to the printer's API to issue a print command.
[1411] Step 10:
[1412] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[1413] Step 11:
[1414] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[1415] Step 12:
[1416] The device triggers a display action, and the image data is sent to the server again.
[1417] Step 13:
[1418] The server sends image data to the TV's API and issues instructions for display.
[1419] Step 14:
[1420] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[1421] Step 15:
[1422] After the server has finished generating, printing, and displaying images, it provides users with a web application containing detailed product information pages that incorporate AI technology.
[1423] Step 16:
[1424] Users can view product information and perform actions to request additional information as needed.
[1425] Step 17:
[1426] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1427] (Example 1)
[1428] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1429] Traditional image generation systems have been criticized for their complex user experience, as the process of generating, displaying, and printing images is cumbersome. Furthermore, they lack mechanisms for quickly providing product information related to the generated images, missing opportunities to increase user purchasing intent. Additionally, there are insufficient methods for users to easily ask questions about products and receive answers.
[1430] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1431] In this invention, the server includes means for hosting a web application accessible by a user terminal, means for receiving generation instructions, means for using a generative artificial intelligence model that generates images based on the generation instructions, means for returning and displaying the generated image data to the user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for displaying information about a product equipped with artificial intelligence technology after the image generation is complete, and means for the user to ask questions about the product through a chatbot and receive responses. This allows the user to easily and quickly perform a series of processes from image generation to display, printing, and product information acquisition, resulting in a high-quality user experience.
[1432] A "user terminal" is a device that a user directly operates to access web applications and the internet.
[1433] A "web application" is a software application that operates via the internet or an intranet and is accessible to users through a web browser.
[1434] "Generation instructions" refer to the text or commands that the user enters to generate an image.
[1435] A "generative artificial intelligence model" is a system that includes machine learning algorithms for generating content such as images based on specified input data.
[1436] A "server" is a computer system on a network that hosts web applications and processes requests from users.
[1437] "Image data" refers to the digital data format of a generated image, which is data that can be displayed and printed.
[1438] A "printing device" is a hardware device that receives digital image data and prints it onto physical paper.
[1439] A "display device" is a device that receives image data and displays it on a screen in high resolution.
[1440] "Product information" refers to data including specifications, reviews, and purchase information related to products equipped with artificial intelligence technology.
[1441] A "chatbot" is conversational software that automatically responds to questions from users.
[1442] Basic configuration
[1443] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. This system includes the following elements:
[1444] 1. Preparing the user interface
[1445] Server: Start a web server (e.g., Apache, Nginx) and host a web application using Python and a framework like Django or Flask. This will allow users to access a webpage with a form to input instructions for image generation.
[1446] 2. Receive instructions for image generation.
[1447] User: Access the web application and enter the content of the image you want to generate as text. For example, use a prompt like "Sunset Landscape". This prompt is submitted when you click the "Generate" button.
[1448] Terminal: Captures the entered text data and sends it to the server.
[1449] 3. Image generation using AI
[1450] Server: Analyzes the received text data and sends it as a prompt to a generative AI model (e.g., OpenAI's DALL-E or GPT-4). Examples of prompts include "sunset landscape" and "picture of a cat."
[1451] Server: Receives image data generated by the generation AI model based on instructions and verifies its format (e.g., PNG, JPEG).
[1452] 4. Displaying the generated image
[1453] Server: Sends the generated image data back to the user's terminal.
[1454] Terminal: Renders received image data on a web browser and displays it to the user.
[1455] 5. Prepare and execute the print job.
[1456] User: Review the displayed image and click the "Print" button on the web application.
[1457] Terminal: Triggers a print action and sends image data to the server.
[1458] Server: Sends image data to the printer's API (e.g., HP's Printer API) and issues a print command.
[1459] Printer: Prints high-quality images using the received image data.
[1460] 6. Preparing and executing the TV display.
[1461] User: Check the displayed image and click the "Display on TV" button.
[1462] Terminal: Triggers a display action and sends image data to the server.
[1463] Server: Sends image data to the TV's API (e.g., Chromecast API) and instructs it to display the image.
[1464] Television: Displays images in high resolution.
[1465] 7. Display of information on AI-equipped products
[1466] Server: After the user completes operations such as image generation, printing, and display, the server provides a page in the web application that displays information about the product, which incorporates artificial intelligence technology. The information page includes detailed product specifications, user reviews, and purchase information.
[1467] Terminal: Displays product information on a webpage, allowing users to view details.
[1468] 8. Considering purchasing a Pixel product
[1469] User: Browse the displayed AI product information page and review the details. Enter questions into the chatbot as needed.
[1470] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[1471] Specific examples of operation
[1472] Example 1: Creating and printing landscape photographs
[1473] User: Enters "I want a landscape photo" into the web application and clicks the generate button.
[1474] Server: Sends instructions as prompts to the AI model generating the data (e.g., DALL-E) and receives the generated landscape photos.
[1475] Terminal: The landscape photo is displayed in a web browser, and the user clicks the print button after confirming it.
[1476] Server: Sends image data to the printer's API and issues a print command.
[1477] Printer: Prints landscape photos in high quality.
[1478] Example 2: Generating a picture of a cat and displaying it on a TV.
[1479] User: Enters "I want to generate a picture of a cat" into the web application and clicks the generate button.
[1480] Server: Sends instructions as prompts to the AI model that generates the data, and receives the generated picture of a cat.
[1481] Device: A picture of a cat is displayed on the webpage, and the user clicks the "Show on TV" button after confirming it.
[1482] Server: Sends image data to the TV's API and issues display instructions.
[1483] Television: Displays a picture of a cat in high resolution.
[1484] Example 3: Information acquisition for AI-powered products
[1485] User: After completing the image generation experience, view the displayed AI product information page to learn more details.
[1486] Server: Provides detailed product specifications, user reviews, and purchase information, and answers questions via chatbot functionality.
[1487] User: Obtain information to help with purchase considerations and resolve questions via chatbots or web pages as needed.
[1488] This invention allows users to easily and quickly perform a series of processes from image generation to display, printing, and acquisition of product information. The system aims to provide a high-quality user experience and enhance user purchasing intent.
[1489] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1490] Program processing steps
[1491] Step 1:
[1492] User: Accesses the web application and enters instructions for image generation. Specifically, the user enters prompt text such as "Sunset Landscape" or "Picture of a Cat" into the text input field and clicks the "Generate" button.
[1493] Input: Prompt text (e.g., "Sunset scenery")
[1494] Output: Text data sent by user interaction
[1495] Step 2:
[1496] Terminal: Sends the entered text data to the server. Specifically, when the user clicks the "Generate" button, that text data is sent to the server as an HTTP request.
[1497] Input: Text data entered by the user
[1498] Output: Text data sent to the server
[1499] Step 3:
[1500] Server: Analyzes the received text data and sends it as a prompt to the generative AI model. Natural language processing (NLP) techniques are used for analysis, and API requests are generated to the generative AI model (e.g., OpenAI's DALL-E or GPT-4).
[1501] Input: Text data received from the device
[1502] Output: API request to the generated AI model
[1503] Step 4:
[1504] Generative AI model: Generates images based on text data. The generated image data is returned to the server.
[1505] Input: Prompt message sent from the server (API request)
[1506] Output: Generated image data
[1507] Step 5:
[1508] Server: Retrieves image data received from the generated AI model and sends it back to the user's terminal. Here, it checks the image data format (e.g., PNG, JPEG) and encodes it in the appropriate format.
[1509] Input: Image data returned from a generative AI model
[1510] Output: Image data to be sent to the user's terminal.
[1511] Step 6:
[1512] Terminal: Renders received image data on a web browser and displays it to the user.
[1513] Input: Image data returned from the server
[1514] Output: Image displayed in the web browser
[1515] Step 7:
[1516] User: Review the generated image and click the "Print" or "Display on TV" button on the web application.
[1517] Input: Image data displayed in a web browser
[1518] Output: User instructions for printing or displaying.
[1519] Step 8:
[1520] Terminal: Triggers user actions (print or view) and sends image data to the server.
[1521] Input: User's print or display instructions
[1522] Output: Image data sent to the server
[1523] Step 9:
[1524] Server: Sends image data to the corresponding device's API to instruct it to print or display. For printing, it sends the data to the printer's API (e.g., HP's Printer API), and for display, it sends the data to the TV's API (e.g., Chromecast API).
[1525] Input: Image data sent from the device
[1526] Output: Image data sent to a printer or television.
[1527] Step 10:
[1528] Printer or television: Receives image data and displays or prints it on the respective device. A printer prints the image on paper, while a television displays the image in high resolution.
[1529] Input: Image data sent from the server
[1530] Output: Printed image or displayed image
[1531] Step 11:
[1532] Server: After the user completes the image generation, printing, and display operations, the server provides a page in the web application that displays product information powered by artificial intelligence technology.
[1533] Input: User operation completion information
[1534] Output: Product information display page
[1535] Step 12:
[1536] User: View information about AI-powered products on a webpage and check the details. Enter questions into the chatbot as needed.
[1537] Input: Viewing product information page
[1538] Output: Question sent to the chatbot
[1539] Step 13:
[1540] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[1541] Input: Question from a user
[1542] Output: Chatbot or human response
[1543] (Application Example 1)
[1544] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1545] Conventional image generation systems had a cumbersome process for users to apply their generated images to specific products and lacked intuitive preview functionality. As a result, users could not visually confirm how the generated images would appear on the actual product, making it difficult to achieve a satisfactory purchasing experience. Furthermore, the means of transmitting the generated images to various output devices for appropriate display and printing were limited.
[1546] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1547] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for applying the generated images to a specific product, means for displaying the result of applying the generated images to a specific product as a three-dimensional preview, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This allows the user to intuitively apply the generated images to a specific product and confirm them in a three-dimensional preview. Furthermore, the generated images can be easily displayed and printed on various output devices, improving the user's purchasing experience.
[1548] A "means for receiving generation instructions" refers to a device or system that receives input from a user to instruct image generation.
[1549] "Means using a generative artificial intelligence model that generates images based on generation instructions" refers to devices or systems that utilize artificial intelligence technology to generate images based on user instructions.
[1550] "Means for displaying generated images on a user's terminal" refers to devices or systems for displaying generated image data on a user's terminal.
[1551] "Means for transmitting generated image data to a printing device and performing printing" refers to a device or system for transmitting generated image data to a printing device and printing the image using the printing device.
[1552] "Means for transmitting generated image data to a display device and displaying it" refers to a device or system for transmitting generated image data to a display device and displaying the image on the display device.
[1553] "Means for applying a generated image to a specific product" refers to a device or system for applying a generated image to a desired product.
[1554] "Means for displaying the result of applying a generated image to a specific product as a three-dimensional preview" refers to a device or system for applying a generated image to a specific product and visualizing and displaying the result in three dimensions.
[1555] "Means for displaying information about a product equipped with artificial intelligence technology after the generation of the aforementioned image is completed" refers to a device or system for providing information about a product equipped with artificial intelligence technology after the generation process is completed.
[1556] The following describes the specific system configuration and program processing for implementing this invention. In particular, it describes the detailed process of generating images using a generative AI model and applying them to a product.
[1557] Basic System Configuration
[1558] 1. Preparing the user interface
[1559] Server: Starts a web server and hosts a web application for users to access. This web application includes a form for entering instructions for image generation.
[1560] 2. Receive instructions for image generation.
[1561] User: Access the web application and enter the content of the image you want to generate as text. Submit this text instruction by clicking the "Generate" button.
[1562] 3. Image generation using AI
[1563] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "Colorful geometric patterns" becomes an instruction to generate an image based on its content.
[1564] Server: The generative artificial intelligence model generates images and sends the results back to the server.
[1565] 4. Displaying the generated image
[1566] Server: Sends the received image data back to the user's terminal.
[1567] Terminal: Receives image data and displays it to the user on the web application.
[1568] 5. Application of images to products
[1569] User: Review the generated image and select the option to apply it to a specific product (e.g., T-shirt, cup, etc.).
[1570] Server: Uses image data to place and apply images to specified products.
[1571] Terminal: Displays a 3D preview of the generated product for the user to review.
[1572] 6. Prepare and execute printing.
[1573] User: After reviewing the generated image, decide to order the product and click the "Print" button.
[1574] Server: Receives print instructions and sends the generated image data to the printing device.
[1575] Printing device: Follow the instructions and print the generated image onto the product.
[1576] 7. Preparing and executing the TV display.
[1577] User: After reviewing the generated image, click the "Display on TV" button.
[1578] Server: Triggers a display action and sends image data to the display device.
[1579] Display device: Displays images in high resolution.
[1580] 8. Display of information on AI-equipped products
[1581] Server: After image generation, printing, and display operations are complete, the server displays a page on the web application that provides users with detailed information about the product, which incorporates artificial intelligence technology.
[1582] Hardware and software to be used
[1583] hardware
[1584] User's smartphone
[1585] Server (e.g., cloud service)
[1586] printing device
[1587] Display device (home television)
[1588] software
[1589] Frontend: React
[1590] Backend: Flask (Python)
[1591] Image generation: OpenAI API
[1592] Specific example
[1593] Example of a prompt
[1594] "Colorful geometric patterns"
[1595] "Spacescape"
[1596] "Cute cat illustration"
[1597] For example, a user accesses a web application, enters "colorful geometric patterns," and clicks the "Generate" button. The server, upon receiving this instruction, uses a generative AI model to generate an image based on the specified content. The generated image is first displayed on the user's terminal. Next, the user applies the image to a specific product and checks it in a 3D preview. If the user is satisfied with this preview, they can send the image to a printing device to print it on the product. The generated image can also be displayed on a home television. Finally, after the image generation experience, the user is provided with detailed information about products equipped with artificial intelligence technology.
[1598] In summary, this invention provides a system that allows users to intuitively generate images and quickly apply them to products for visualization. Furthermore, user satisfaction can be enhanced by displaying or printing the generated images on various output devices.
[1599] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1600] Step 1:
[1601] The user enters a prompt.
[1602] The user accesses the web application and enters a text prompt for image generation (e.g., "Colorful geometric pattern"). Once the input is complete, they click the "Generate" button.
[1603] Input: Text prompt (e.g., "Colorful geometric patterns")
[1604] Output: Generation instructions
[1605] Specific operation: The user enters a prompt message in the text input field and clicks a button, which sends the instruction to the server.
[1606] Step 2:
[1607] The server receives the generation instruction.
[1608] The server receives the prompt message sent by the user and parses it.
[1609] Input: Generation Instructions
[1610] Output: Parsed prompt message
[1611] Specific operation: The server receives an HTTP request and extracts the prompt text.
[1612] Step 3:
[1613] The server sends a request to the generated AI model.
[1614] Based on the parsed prompt text, the server sends an image generation request to the AI model.
[1615] Input: Parsed prompt message
[1616] Output: Request for generated AI model
[1617] Specific operation: The server generates a prompt message and sends it to the AI model's API in the appropriate format.
[1618] Step 4:
[1619] The generative AI model generates images.
[1620] The generative AI model generates an image based on the prompt text and sends the result back to the server.
[1621] Input: Request for generated AI model
[1622] Output: Generated image data
[1623] Specific operation: The generative AI model performs internal processing, generates an image based on the prompt text, and sends that data back to the server.
[1624] Step 5:
[1625] The server receives the generated image and sends it to the user's terminal.
[1626] The server receives the generated image data and sends it back to the user's terminal.
[1627] Input: Generated image data
[1628] Output: Sending image data to the user terminal
[1629] Specific operation: The server sends the received image data to the user's terminal as an HTTP response.
[1630] Step 6:
[1631] Display images on the user's terminal.
[1632] The user terminal displays the received image data on the web application.
[1633] Input: Image data sent to the user terminal
[1634] Output: Image display on a web application
[1635] Specific operation: The user's device browser renders the image data and displays it on the screen.
[1636] Step 7:
[1637] Users apply images to the product.
[1638] The user reviews the generated image and selects the option to apply it to a specific product (e.g., a T-shirt, a cup).
[1639] Input: Generated image and product selection
[1640] Output: Image applied to the product
[1641] Specific operation: The user clicks a product selection option and sends an apply request to the server.
[1642] Step 8:
[1643] The server applies the image to the product.
[1644] The server applies the generated image to a specific product and generates the corresponding data.
[1645] Input: Generated image and product selection information
[1646] Output: Image data applied to the product
[1647] Specific operation: The server places images into the specified product template and generates the applied data.
[1648] Step 9:
[1649] The server sends a 3D preview to the user terminal.
[1650] The server sends the generated product data to the user terminal in a three-dimensional preview format.
[1651] Input: Image data applied to the product
[1652] Output: Sending a 3D preview to the user's terminal
[1653] Specific operation: The server generates data for the 3D preview and sends it to the user's terminal as an HTTP response.
[1654] Step 10:
[1655] The user terminal displays a 3D preview.
[1656] The user terminal displays the received 3D preview on the web application.
[1657] Input: Data for 3D preview
[1658] Output: Three-dimensional preview display on screen
[1659] Specific operation: The user's browser renders a 3D preview and displays it on the screen.
[1660] Step 11:
[1661] The user issues a print command.
[1662] The user reviews the 3D preview, and if satisfied, clicks the "Print" button to initiate printing.
[1663] Input: User's print instructions
[1664] Output: Print request to server
[1665] Specific action: The user clicks a button and sends a print command to the server.
[1666] Step 12:
[1667] The server sends image data to the printer.
[1668] The server receives the print command and sends the generated image data to the printing device.
[1669] Input: User's print instructions and image data.
[1670] Output: Print instructions to the printer
[1671] Specific operation: The server sends image data to the printer's API and issues a print command.
[1672] Step 13:
[1673] The printing device prints the image.
[1674] The printing device prints an image onto the specified product based on the received image data.
[1675] Input: Print instructions to the printer and image data.
[1676] Output: Printed product
[1677] Specific operation: The printing device prints an image onto the product based on the data it receives.
[1678] Step 14:
[1679] The user displays an image on the display device.
[1680] The user clicks the "Display on TV" button to display the generated image on a display device such as a home television.
[1681] Input: User's display instructions
[1682] Output: Display request to the server
[1683] Specific operation: The user clicks a button and sends a display instruction to the server.
[1684] Step 15:
[1685] The server sends image data to the display device.
[1686] The server receives the display instruction and sends the generated image data to the display device.
[1687] Input: User display instructions and image data
[1688] Output: Display instructions for the display device.
[1689] Specific operation: The server sends image data to the display device's API and issues a display instruction.
[1690] Step 16:
[1691] The display device displays an image.
[1692] The display device displays the image in high resolution based on the received image data.
[1693] Input: Display instructions and image data for the display device.
[1694] Output: Displayed image
[1695] Specific operation: The display device displays an image on the screen based on the data it receives.
[1696] Step 17:
[1697] The server displays information about products equipped with AI technology.
[1698] After the server completes image generation, printing, and display operations, it displays a page on the web application that provides users with detailed information about the AI-powered product.
[1699] Input: User's operation complete
[1700] Output: Information on products equipped with AI technology
[1701] Specific action: The server updates the content of the webpage and provides the user with detailed information.
[1702] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1703] This invention relates to a generative AI system that combines an emotion engine to recognize user emotions. This system receives image generation instructions from the user, generates images using a generative artificial intelligence model, and outputs the generated images to a printer or display device. Furthermore, it displays information about products equipped with artificial intelligence technology after image generation. The system incorporates an emotion engine that recognizes the user's emotional state and adjusts generation instructions based on that state.
[1704] Basic configuration
[1705] 1. Preparing the user interface
[1706] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[1707] 2. Initiation of emotion recognition
[1708] The user accesses the web application and clicks a specific button to initiate emotion recognition.
[1709] The device collects user interaction and input data and sends it to the emotion engine.
[1710] 3. Analysis of emotional state
[1711] The server uses an emotion engine to analyze the user's emotional state. For example, it identifies emotions based on the user's input speed, input content, and facial expression data (if acquired using a camera).
[1712] 4. Receive instructions for image generation.
[1713] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[1714] The emotion engine takes the user's emotional state into consideration and makes adaptive adjustments to the generation instructions.
[1715] The terminal sends the adjusted generation instruction data to the server.
[1716] 5. Image generation using AI
[1717] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, based on the input text "Sunset Landscape," it generates an image that contains emotionally positive elements.
[1718] 6. Displaying the generated image
[1719] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user's terminal.
[1720] The device receives image data and displays it to the user on the web application. The user then views the generated image on the screen.
[1721] 7. Prepare and execute the print job.
[1722] The user reviews the generated image and clicks the "Print" button on the web application.
[1723] The terminal triggers a print action and resends the image data to the server.
[1724] The server receives the image data and sends it to the printer's API to issue a print command.
[1725] The printer prints the image in high quality according to the instructions. The user receives the printed image.
[1726] 8. Preparing and executing the TV display.
[1727] After the user reviews the generated image, they click the "Display on TV" button.
[1728] The device triggers a display action, and the image data is sent to the server again.
[1729] The server sends image data to the TV's API and issues instructions for display.
[1730] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[1731] 9. Information display for AI-equipped products
[1732] After the server has finished generating, printing, and displaying images, it provides users with detailed information about the AI-powered product via a web application.
[1733] The device displays the product information page to the user.
[1734] 10. Considering purchasing a Pixel product
[1735] Users view product information and, if interested, take actions to request additional information.
[1736] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1737] Specific example
[1738] Example 1: Emotion-driven landscape photography and printing
[1739] The user enters "I want a landscape photo" and clicks the "Generate" button.
[1740] The emotion engine analyzes user input and interaction data to identify positive emotional states.
[1741] The server sends an instruction to the AI model to generate a "positive sunset landscape."
[1742] The server receives the generated landscape photo data and sends it to the terminal.
[1743] The device displays landscape photos it has received.
[1744] After the user reviews the image, they click the "Print" button.
[1745] The server sends the image data to the printer and starts printing.
[1746] The printer prints high-quality landscape photographs.
[1747] Example 2: Generation and display of cat images based on emotions
[1748] The user enters "I want to generate a picture of a cat" and clicks the "Generate" button.
[1749] The emotion engine analyzes the user's interaction data and determines that the user is slightly tired.
[1750] The server sends an instruction to the AI model to generate a "relaxing picture of a cat."
[1751] The server receives the generated cat image data and sends it to the terminal.
[1752] The device displays the image of a cat it received.
[1753] After the user reviews the image, they click the "Display on TV" button.
[1754] The server sends the image data to the television and begins displaying it.
[1755] The TV displays a picture of a cat in high resolution.
[1756] Example 3: Information retrieval for AI-powered products based on emotions
[1757] After the user completes the image generation experience, they can view information pages about AI-powered products on the web application.
[1758] The emotion engine recognizes the user's level of excitement and prioritizes displaying highly relevant product information.
[1759] The server provides detailed product information, and a chatbot answers questions.
[1760] Users can obtain information to consider purchasing and inquire about details as needed.
[1761] In this way, a system is realized that takes user emotions into consideration and provides a consistent service from image generation using generative AI to output methods and product information acquisition. By incorporating an emotion engine, the user experience can be further personalized and satisfaction can be increased.
[1762] The following describes the processing flow.
[1763] Step 1:
[1764] The server starts up a web server and hosts a web application for users to access. This allows users to access a form to enter instructions for image generation.
[1765] Step 2:
[1766] The user accesses a web application and clicks a specific button to initiate emotion recognition. This action causes the emotion engine to begin collecting data about the user.
[1767] Step 3:
[1768] The device collects user interaction data (such as input speed, input content, and in some cases, facial recognition data from the webcam) and sends it to the emotion engine.
[1769] Step 4:
[1770] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the collected data to determine whether the user is in a positive, negative, relaxed, or other emotional state.
[1771] Step 5:
[1772] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[1773] Step 6:
[1774] The emotion engine considers the user's emotional state and makes adaptive adjustments to the "generate" instructions. For example, if the user is in a positive state, it adds instructions to include vibrant colors and bright scenes in the generated image.
[1775] Step 7:
[1776] The terminal sends the adjusted generation instruction data to the server. The transmitted data includes text as well as adjustment information for the emotion engine.
[1777] Step 8:
[1778] The server analyzes the received text data and sentiment adjustment data, and sends an image generation request to the generative artificial intelligence model. For example, when providing the generative AI model with the text "sunset landscape," instructions are also given to include positive elements.
[1779] Step 9:
[1780] The server receives the generated image data returned from the generated artificial intelligence model and sends it to the user's terminal. The image data is usually encoded in a format such as Base64.
[1781] Step 10:
[1782] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[1783] Step 11:
[1784] After the user reviews the generated image, they click the "Print" button on the web application.
[1785] Step 12:
[1786] The terminal triggers a print action and resends the image data to the server.
[1787] Step 13:
[1788] The server receives the image data and sends it to the printer's API to issue a print command.
[1789] Step 14:
[1790] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[1791] Step 15:
[1792] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[1793] Step 16:
[1794] The device triggers a display action, and the image data is sent to the server again.
[1795] Step 17:
[1796] The server sends image data to the TV's API and issues instructions for display.
[1797] Step 18:
[1798] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[1799] Step 19:
[1800] After the server completes the image generation, printing, and display operations, it provides users with detailed information about the product, which incorporates artificial intelligence technology, through a web application.
[1801] Step 20:
[1802] Users view the displayed product information and, if interested, take action to request additional information.
[1803] Step 21:
[1804] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[1805] (Example 2)
[1806] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1807] Conventional image generation systems often fail to consider the user's emotional state when issuing generation instructions, resulting in generated images that do not always match the user's expectations or feelings. Furthermore, the limited output formats and methods for generated images posed a challenge, hindering user convenience.
[1808] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1809] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for adjusting the generation instructions using an emotion engine that recognizes the user's emotional state, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This makes it possible to generate images that take the user's emotional state into consideration, thereby increasing user satisfaction. In addition, the generated images can be output to a variety of devices, improving user convenience.
[1810] "Generation instructions" are data entered by the user to specify the content and conditions of the image that should be generated.
[1811] A "generative artificial intelligence model" is an artificial intelligence algorithm or system that generates images based on instructions or data input by a user.
[1812] A "user terminal" refers to a device such as a computer or smartphone that a user operates, and which provides an interface for image generation through a web application.
[1813] A "printing device" is a hardware device used to physically print generated images, and primarily refers to a printer.
[1814] A "display device" is a hardware device used to visually present generated images to a user, and includes televisions and monitors.
[1815] An "emotion engine" is software or a system that analyzes user interaction data and input to identify emotional states and adjusts generation instructions based on those results.
[1816] "Information about products equipped with artificial intelligence technology" refers to detailed information about products incorporating artificial intelligence technology, which is provided to the user after the display of the generated image is complete.
[1817] This invention relates to an image generation system that incorporates an emotion engine to recognize user emotions. The system receives image generation instructions from the user and generates images using a generation artificial intelligence model. The generated images are then output to a printer or display device, and information about the product, which incorporates artificial intelligence technology, is displayed after image generation.
[1818] composition
[1819] The system of the present invention has the following components.
[1820] 1. Preparing the user interface
[1821] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is implemented using HTML / CSS / JavaScript.
[1822] 2. Initiation of emotion recognition
[1823] The user clicks the "Start Emotion Recognition" button in the web application. The device captures the click event and prepares to send interaction data to the emotion engine. This is done using services such as Microsoft Azure's Emotion API.
[1824] 3. Analysis of emotional state
[1825] The server sends the received interaction data to the emotion engine, which then analyzes the user's emotional state. For example, it analyzes input content, input speed, and facial expression data captured by the camera.
[1826] 4. Receive instructions for image generation.
[1827] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the emotional state and makes adaptive adjustments to the generation instructions. The device then sends the adjusted generation instruction data to the server.
[1828] 5. Image generation using AI
[1829] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[1830] 6. Displaying the generated image
[1831] The server receives the generated image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[1832] 7. Prepare and execute the print job.
[1833] The user reviews the generated image and clicks the "Print" button. The device resends the image data to the server. The server sends the data to the printer's API (such as Google Cloud Print) and prints the image.
[1834] 8. Preparing and executing the TV display.
[1835] After the user reviews the generated image, they click the "Display on TV" button. The device sends the image data to the server, which then sends it to the TV's API (such as Google Cast or Samsung Smart View API) for display.
[1836] 9. Information display for AI-equipped products
[1837] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[1838] 10. Considering purchasing a Pixel product
[1839] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[1840] Specific example
[1841] Example 1: Creating and printing landscape photographs based on emotions.
[1842] User: Enter "I want a landscape photo" and click the "Generate" button.
[1843] Emotion Engine: Analyzes user input and emotional states to identify positive emotional states.
[1844] Server: Sends the instruction "positive sunset scenery" to the generating artificial intelligence model.
[1845] Server: Receives the generated image data and sends it to the terminal.
[1846] Terminal: Receives and displays image data.
[1847] User: Click the "Print" button.
[1848] Server: Sends image data to the printer API and starts printing.
[1849] Printer: Prints high-quality landscape photos.
[1850] Example 2: Generation of cat images based on emotions and display on television
[1851] User: Enter "I want to generate a picture of a cat" and click the "Generate" button.
[1852] Emotion Engine: Analyzes user interaction data to recognize fatigue levels.
[1853] Server: Instructs the artificial intelligence model to generate a "relaxing picture of a cat".
[1854] Server: Receives the generated cat image data and sends it to the terminal.
[1855] Terminal: Receives and displays image data.
[1856] User: Click the "Display on TV" button.
[1857] Server: Sends image data to the TV API and starts displaying it on the TV.
[1858] Television: Displays a picture of a cat in high resolution.
[1859] Example 3: Information acquisition for AI-powered products based on emotions
[1860] User: After completing the image generation experience, they browse the information page for AI-powered products in the web application.
[1861] Emotion Engine: Recognizes the user's emotional state and prioritizes displaying highly relevant product information.
[1862] Server: Provides detailed product information and answers questions via chatbot.
[1863] User: Obtain information to consider purchasing and inquire for further details as needed.
[1864] In this way, by incorporating an emotion engine, we can personalize the user experience and create a system that enhances satisfaction. The collaboration between the generative AI model and the emotion engine enables the generation and output of appropriate images that take into account the user's emotional state.
[1865] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1866] Step 1: Preparing the User Interface
[1867] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is constructed using HTML, CSS, and JavaScript. Users access it through a web browser and are presented with a form to input instructions for image generation.
[1868] Input: User access request
[1869] Output: Display of a form for entering image generation instructions.
[1870] Specific operation: The server starts up, and the web application waits for user requests. When a user accesses the web application, an instruction input form is displayed.
[1871] Step 2: Initiating Emotion Recognition
[1872] The user clicks the "Start emotion recognition" button in the web application. The device captures the click event and collects user interaction data (facial expressions, typing speed, text input, etc.).
[1873] Input: User click event
[1874] Output: Start of interaction data collection
[1875] Specific actions: The user clicks a button, the device detects the click event and turns on the camera. The device prepares to collect facial expression data and text input data.
[1876] Step 3: Analyze your emotional state
[1877] The server sends interaction data from the terminal to an emotion engine (for example, Microsoft Azure's Emotion API) to analyze the user's emotional state.
[1878] Input: Interaction data from the device
[1879] Output: Sentiment analysis results from the emotion engine
[1880] Specific operation: The server receives data sent from the terminal and sends it to the emotion engine. The emotion engine identifies the emotional state and returns the result to the server.
[1881] Step 4: Receive instructions for image generation.
[1882] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the user's emotional state and adaptively adjusts the generation instructions. The device then sends the adjusted generation instruction data to the server.
[1883] Input: User's image generation instructions
[1884] Output: Adjusted generation instruction data
[1885] Specific operation: The user enters the image content as text and clicks the "Generate" button. The device sends this input data to the server, which then adjusts the instructions via the emotion engine.
[1886] Step 5: Image generation using AI
[1887] The server analyzes the received text data and sends an image generation request to a generative artificial intelligence model (such as OpenAI's DALL-E or Google's Imagen).
[1888] Input: Adjusted generation instruction data
[1889] Output: Generated image data
[1890] Specific operation: The server analyzes the text data and sends it to the generative AI model as a prompt. The generative AI model generates an image and sends it back to the server.
[1891] Step 6: Displaying the generated image
[1892] The server receives the generated image data returned from the generated AI model and sends it to the terminal. The terminal receives the image data and displays it to the user on the web application.
[1893] Input: Generated image data
[1894] Output: Image display on a web application
[1895] Specific operation: The server receives image data and sends it to the terminal. The terminal receives the image data and displays it in the web application.
[1896] Step 7: Prepare and execute print
[1897] The user reviews the generated image and clicks the "Print" button. The device triggers the print action and resends the image data to the server. The server sends the image data to the printer API and starts printing.
[1898] Input: User's print instructions
[1899] Output: Printed image
[1900] Specific operation: The user reviews the image and clicks the "Print" button. The device sends the image data to the server, and the server sends a print command to the printer API. The printer prints the image in high quality.
[1901] Step 8: Prepare and run the TV display.
[1902] After the user reviews the generated image, they click the "Display on TV" button. The device triggers the display action and sends the image data to the server again. The server sends the image data to the TV's API and starts displaying it.
[1903] Input: User's display instructions
[1904] Output: Image displayed on the TV
[1905] Specific steps: The user views the image and clicks the "Display on TV" button. The device sends the image data to the server, and the server sends a display command to the TV API. The TV displays the image in high resolution.
[1906] Step 9: Displaying information about AI-powered products
[1907] After the server has finished generating, printing, and displaying the image, the web application provides the user with detailed information about the AI-powered product. The device then displays the product information page.
[1908] Input: State after operation is complete
[1909] Output: Product Information Page
[1910] Specific operation: The server adds product information to the web application, and the terminal displays the product information page to the user.
[1911] Step 10: Consider purchasing a Pixel product
[1912] If a user views product information and becomes interested, they enter a question into the chatbot. The server automatically responds to the user's question through the chatbot and, if necessary, transfers the user to a human representative.
[1913] Input: User's question
[1914] Output: Automated chatbot response and transfer to a human resource representative.
[1915] Specific operation: The user views product information and enters a question into the chatbot. The server provides an automated response and, if necessary, transfers the user to a human representative.
[1916] (Application Example 2)
[1917] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1918] Conventional image generation systems do not provide personalized image generation or content recommendations based on the user's emotional state. This results in a limited user experience and potentially lower satisfaction. Furthermore, there is a need to integrate emotion recognition technology into image generation and content recommendations to provide services that are more emotionally resonant.
[1919] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1920] In this invention, the server includes means for receiving generation instructions, means for using a generative artificial intelligence model to generate images based on the generation instructions, means for analyzing the user's emotional state using emotion recognition means, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, and means for displaying information about products equipped with artificial intelligence technology after the image generation is complete. This enables personalized image generation and content recommendations according to the user's emotional state.
[1921] A "generation instruction" refers to a request regarding images or content that the user will generate.
[1922] A "generative artificial intelligence model" is a part of an artificial intelligence system that includes algorithms and databases for generating images and content based on user instructions and data.
[1923] "Emotion recognition means" refers to a part of a system that includes hardware and software for analyzing the user's emotional state from facial expressions, tone of voice, input speed, etc.
[1924] "Generated image data" refers to the digital representation of image information generated by a generative artificial intelligence model.
[1925] A "user terminal" is a device used to display or print generated images and recommended content.
[1926] A "printing device" is a machine used to print generated image data into a physical format.
[1927] A "display device" is a screen or monitor used to visually display generated image data or recommended content.
[1928] "Products equipped with artificial intelligence technology" refer to goods and services that provide functions using artificial intelligence technologies such as image generation and emotion recognition.
[1929] "Personalized content" refers to content that is customized for a specific user based on their emotional state and preferences.
[1930] The system for implementing this invention is an emotion recognition and content generation system that recognizes the user's emotions and generates and recommends personalized content accordingly.
[1931] This system includes a means of receiving generation instructions from users. A server hosts a web application with a user interface, from which users input instructions for generating images and content. These instructions include the content of the images and content desired by the user.
[1932] The server uses a generative artificial intelligence model to generate images and content based on generation instructions. This model has the ability to generate high-quality images and content based on input such as text prompts.
[1933] Next, the emotion recognition system analyzes the user's emotional state. This system includes a camera and microphone to capture the user's facial expressions, voice tone, input speed, etc., and software to analyze the input data. EmotionEngine, using TensorFlow, identifies the emotional state.
[1934] The server displays the generated image data and content on the user's terminal. The user's terminal is a compatible device such as a smartphone, smart glasses, or head-mounted display.
[1935] Furthermore, it includes means for sending the generated image data to a printing device and printing it in physical form. To achieve this, the server uses the printer's API to send the image data and issue print commands.
[1936] In addition, the system includes means for transmitting the generated image data to a display device and displaying it on a large screen or high-resolution display.
[1937] Finally, the server displays information about the AI-powered product after image generation is complete. This allows users to obtain detailed information and purchase options for the AI-powered product associated with the generated content.
[1938] For example, if a user feels the need to relax, the emotion recognition system analyzes the user's emotional state and generates and recommends a playlist of relaxing music based on that analysis. Examples of generation instructions could be text inputs such as, "Generate a music playlist suitable for when the user wants to relax," or "Recommend entertainment videos that match the user's current mood."
[1939] This system allows users to receive personalized images and content adapted to their emotional state in real time, resulting in a more satisfying user experience.
[1940] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1941] Step 1:
[1942] The user accesses the web application and enters generation instructions.
[1943] Input: User text input (e.g., "I want to generate relaxing images")
[1944] Output: Generation instruction data
[1945] Specific action: The user enters text into a designated form on the web application and clicks the "Generate" button.
[1946] Step 2:
[1947] The server begins emotion recognition.
[1948] Input: User's facial expression data, voice data, interaction data such as input speed, etc.
[1949] Output: Emotional state data (e.g., "I want to relax")
[1950] Specific operation: The server uses EmotionEngine to analyze data collected from the camera and microphone to identify the user's emotional state.
[1951] Step 3:
[1952] The server adjusts the generation instructions based on the emotional state.
[1953] Input: Emotional state data, generation instruction data
[1954] Output: Adjusted generation instruction data (e.g., "Relaxing sunset landscape image")
[1955] Specific operation: Based on emotional state data, personalized elements are added to the generation instruction data.
[1956] Step 4:
[1957] The server sends generation instructions to the artificial intelligence model, which then generates the image.
[1958] Input: Adjusted generation instruction data
[1959] Output: Generated image data
[1960] Specific operation: The server sends a prompt message to the generation AI model and generates the specified image (e.g., "generate a relaxing sunset landscape").
[1961] Step 5:
[1962] The server sends the generated image data to the user's terminal for display.
[1963] Input: Generated image data
[1964] Output: Image displayed on the user's terminal
[1965] Specific operation: The server encodes the generated image data and sends it to the user's terminal. The user's terminal decodes the image data and displays it on the web application.
[1966] Step 6:
[1967] The user prints the generated image data.
[1968] Input: Generated image data, user's print instructions
[1969] Output: Printed image
[1970] Specific operation: When the user clicks the "Print" button, the server sends the image data to the printer's API and performs high-quality printing.
[1971] Step 7:
[1972] The user sends generated image data to a display device, which then displays it on a large screen.
[1973] Input: Generated image data, user display instructions
[1974] Output: Image displayed on the display device
[1975] Specific operation: When the user clicks the "Display on TV" button, the server sends the image data to the display device's API and displays it on the large screen in high resolution.
[1976] Step 8:
[1977] After the server finishes generating the image, it displays information about the product, which incorporates artificial intelligence technology, on the user's terminal.
[1978] Input: Generated image data, related product information
[1979] Output: Product information displayed on the user terminal
[1980] Specific operation: The server retrieves relevant AI product information from the generated image data and displays it on the user's terminal. This information includes detailed product information and purchase links.
[1981] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1982] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1983] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1984] [Fourth Embodiment]
[1985] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1986] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1987] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1988] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1989] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1990] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1991] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1992] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1993] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1994] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1995] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1996] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1997] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1998] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. In the following description, the program processing at each step of the system will be explained in natural language, with specific examples.
[1999] Basic configuration
[2000] 1. Preparing the user interface
[2001] Server: Starts a web server and hosts the web application that users can access. This allows users to access a form to enter instructions for image generation.
[2002] 2. Receive instructions for image generation.
[2003] User: Access the web application and enter the content of the image you want to generate as text. Submit this instruction by clicking the "Generate" button.
[2004] Terminal: Receives the entered text data and sends it to the server.
[2005] 3. Image generation using AI
[2006] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "sunset landscape" would be an instruction to generate an image depicting an evening sky and horizon.
[2007] Server: The generative artificial intelligence model generates images and sends the generation results back to the server.
[2008] 4. Displaying the generated image
[2009] Server: Sends the received image data back to the user's terminal.
[2010] Terminal: Receives image data and displays it to the user on the web application.
[2011] 5. Prepare and execute the print job.
[2012] User: Review the generated image and click the "Print" button on the web application.
[2013] Terminal: Triggers a print action and sends image data to the server.
[2014] Server: Sends the received image data to the printer's API and instructs it to print.
[2015] Printer: Print the image according to the instructions. The user can verify that the generated image is printed in high quality.
[2016] 6. Preparing and executing the TV display.
[2017] User: Review the generated image and click the "Display on TV" button.
[2018] Terminal: Triggers a display action and sends image data to the server.
[2019] Server: Sends the received image data to the TV's API and instructs it to display the image.
[2020] Television: Displays images in high resolution. Users can see that the generated images are displayed on a large screen.
[2021] 7. Display of information on AI-equipped products
[2022] Server: After image generation, printing, and display operations are complete, the server provides the user with a web application page displaying detailed information about the product, which incorporates artificial intelligence technology.
[2023] Terminal: Users can view product information they are interested in and perform actions to obtain more detailed information.
[2024] 8. Considering purchasing a Pixel product
[2025] User: View the provided product information, obtain additional information as needed, and consider purchasing the product.
[2026] Server: Provides automated responses via chatbots and handover functions to human staff in response to user actions.
[2027] Specific example
[2028] Example 1: Creating and printing landscape photographs
[2029] User: Enters "I want a landscape photo" into the web application and clicks the "Generate" button.
[2030] Server: Sends user instructions to the generation AI and receives the generated landscape photo data.
[2031] Terminal: Displays received landscape photos.
[2032] User: After reviewing the image, click the "Print" button.
[2033] Server: Sends image data to the printer and starts printing.
[2034] Printer: Prints high-quality landscape photos.
[2035] Example 2: Generating a picture of a cat and displaying it on a TV.
[2036] User: Type "I want to generate a picture of a cat" and click the "Generate" button.
[2037] Server: Sends user instructions to the generation AI and receives the generated cat image data.
[2038] Terminal: Displays the received picture of a cat.
[2039] User: After reviewing the image, click the "Display on TV" button.
[2040] Server: Sends image data to the television and starts displaying it.
[2041] Television: Displays a picture of a cat in high resolution.
[2042] Example 3: Information acquisition for AI-powered products
[2043] User: After completing the image generation experience, view the displayed information page for AI-powered products.
[2044] Server: Provides detailed product information and answers questions via chatbot.
[2045] User: Obtain information to consider purchasing and inquire for further details as needed.
[2046] As described above, this system utilizes generative AI to generate images and provides a series of processes for outputting those images in various ways. Furthermore, by providing users with information about AI-powered products through the generation experience, it becomes possible to create a new purchasing experience.
[2047] The following describes the processing flow.
[2048] Step 1:
[2049] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[2050] Step 2:
[2051] The user accesses the web application and enters the content of the image they want to generate as text. Once the input is complete, they click the "Generate" button.
[2052] Step 3:
[2053] The terminal receives user input and sends text data to the server. The transmitted data is in JSON format.
[2054] Step 4:
[2055] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, it passes the text "sunset landscape" to the "generative AI".
[2056] Step 5:
[2057] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user terminal. The image data is usually encoded in a format such as Base64.
[2058] Step 6:
[2059] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[2060] Step 7:
[2061] After the user reviews the generated image, they click the "Print" button on the web application.
[2062] Step 8:
[2063] The terminal triggers a print action and resends the image data to the server.
[2064] Step 9:
[2065] The server receives the image data and sends it to the printer's API to issue a print command.
[2066] Step 10:
[2067] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[2068] Step 11:
[2069] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[2070] Step 12:
[2071] The device triggers a display action, and the image data is sent to the server again.
[2072] Step 13:
[2073] The server sends image data to the TV's API and issues instructions for display.
[2074] Step 14:
[2075] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[2076] Step 15:
[2077] After the server has finished generating, printing, and displaying images, it provides users with a web application containing detailed product information pages that incorporate AI technology.
[2078] Step 16:
[2079] Users can view product information and perform actions to request additional information as needed.
[2080] Step 17:
[2081] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[2082] (Example 1)
[2083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2084] Traditional image generation systems have been criticized for their complex user experience, as the process of generating, displaying, and printing images is cumbersome. Furthermore, they lack mechanisms for quickly providing product information related to the generated images, missing opportunities to increase user purchasing intent. Additionally, there are insufficient methods for users to easily ask questions about products and receive answers.
[2085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[2086] In this invention, the server includes means for hosting a web application accessible by a user terminal, means for receiving generation instructions, means for using a generative artificial intelligence model that generates images based on the generation instructions, means for returning and displaying the generated image data to the user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for displaying information about a product equipped with artificial intelligence technology after the image generation is complete, and means for the user to ask questions about the product through a chatbot and receive responses. This allows the user to easily and quickly perform a series of processes from image generation to display, printing, and product information acquisition, resulting in a high-quality user experience.
[2087] A "user terminal" is a device that a user directly operates to access web applications and the internet.
[2088] A "web application" is a software application that operates via the internet or an intranet and is accessible to users through a web browser.
[2089] "Generation instructions" refer to the text or commands that the user enters to generate an image.
[2090] A "generative artificial intelligence model" is a system that includes machine learning algorithms for generating content such as images based on specified input data.
[2091] A "server" is a computer system on a network that hosts web applications and processes requests from users.
[2092] "Image data" refers to the digital data format of a generated image, which is data that can be displayed and printed.
[2093] A "printing device" is a hardware device that receives digital image data and prints it onto physical paper.
[2094] A "display device" is a device that receives image data and displays it on a screen in high resolution.
[2095] "Product information" refers to data including specifications, reviews, and purchase information related to products equipped with artificial intelligence technology.
[2096] A "chatbot" is conversational software that automatically responds to questions from users.
[2097] Basic configuration
[2098] This embodiment of the invention is a system in which a user inputs generation instructions, generates an image using a generation artificial intelligence model, outputs the image to a printer or display device, and further displays information about the product equipped with artificial intelligence technology. This system includes the following elements:
[2099] 1. Preparing the user interface
[2100] Server: Start a web server (e.g., Apache, Nginx) and host a web application using Python and a framework like Django or Flask. This will allow users to access a webpage with a form to input instructions for image generation.
[2101] 2. Receive instructions for image generation.
[2102] User: Access the web application and enter the content of the image you want to generate as text. For example, use a prompt like "Sunset Landscape". This prompt is submitted when you click the "Generate" button.
[2103] Terminal: Captures the entered text data and sends it to the server.
[2104] 3. Image generation using AI
[2105] Server: Analyzes the received text data and sends it as a prompt to a generative AI model (e.g., OpenAI's DALL-E or GPT-4). Examples of prompts include "sunset landscape" and "picture of a cat."
[2106] Server: Receives image data generated by the generation AI model based on instructions and verifies its format (e.g., PNG, JPEG).
[2107] 4. Displaying the generated image
[2108] Server: Sends the generated image data back to the user's terminal.
[2109] Terminal: Renders received image data on a web browser and displays it to the user.
[2110] 5. Prepare and execute the print job.
[2111] User: Review the displayed image and click the "Print" button on the web application.
[2112] Terminal: Triggers a print action and sends image data to the server.
[2113] Server: Sends image data to the printer's API (e.g., HP's Printer API) and issues a print command.
[2114] Printer: Prints high-quality images using the received image data.
[2115] 6. Preparing and executing the TV display.
[2116] User: Check the displayed image and click the "Display on TV" button.
[2117] Terminal: Triggers a display action and sends image data to the server.
[2118] Server: Sends image data to the TV's API (e.g., Chromecast API) and instructs it to display the image.
[2119] Television: Displays images in high resolution.
[2120] 7. Display of information on AI-equipped products
[2121] Server: After the user completes operations such as image generation, printing, and display, the server provides a page in the web application that displays information about the product, which incorporates artificial intelligence technology. The information page includes detailed product specifications, user reviews, and purchase information.
[2122] Terminal: Displays product information on a webpage, allowing users to view details.
[2123] 8. Considering purchasing a Pixel product
[2124] User: Browse the displayed AI product information page and review the details. Enter questions into the chatbot as needed.
[2125] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[2126] Specific examples of operation
[2127] Example 1: Creating and printing landscape photographs
[2128] User: Enters "I want a landscape photo" into the web application and clicks the generate button.
[2129] Server: Sends instructions as prompts to the AI model generating the data (e.g., DALL-E) and receives the generated landscape photos.
[2130] Terminal: The landscape photo is displayed in a web browser, and the user clicks the print button after confirming it.
[2131] Server: Sends image data to the printer's API and issues a print command.
[2132] Printer: Prints landscape photos in high quality.
[2133] Example 2: Generating a picture of a cat and displaying it on a TV.
[2134] User: Enters "I want to generate a picture of a cat" into the web application and clicks the generate button.
[2135] Server: Sends instructions as prompts to the AI model that generates the data, and receives the generated picture of a cat.
[2136] Device: A picture of a cat is displayed on the webpage, and the user clicks the "Show on TV" button after confirming it.
[2137] Server: Sends image data to the TV's API and issues display instructions.
[2138] Television: Displays a picture of a cat in high resolution.
[2139] Example 3: Information acquisition for AI-powered products
[2140] User: After completing the image generation experience, view the displayed AI product information page to learn more details.
[2141] Server: Provides detailed product specifications, user reviews, and purchase information, and answers questions via chatbot functionality.
[2142] User: Obtain information to help with purchase considerations and resolve questions via chatbots or web pages as needed.
[2143] This invention allows users to easily and quickly perform a series of processes from image generation to display, printing, and acquisition of product information. The system aims to provide a high-quality user experience and enhance user purchasing intent.
[2144] The flow of the specific processing in Example 1 will be explained using Figure 11.
[2145] Program processing steps
[2146] Step 1:
[2147] User: Accesses the web application and enters instructions for image generation. Specifically, the user enters prompt text such as "Sunset Landscape" or "Picture of a Cat" into the text input field and clicks the "Generate" button.
[2148] Input: Prompt text (e.g., "Sunset scenery")
[2149] Output: Text data sent by user interaction
[2150] Step 2:
[2151] Terminal: Sends the entered text data to the server. Specifically, when the user clicks the "Generate" button, that text data is sent to the server as an HTTP request.
[2152] Input: Text data entered by the user
[2153] Output: Text data sent to the server
[2154] Step 3:
[2155] Server: Analyzes the received text data and sends it as a prompt to the generative AI model. Natural language processing (NLP) techniques are used for analysis, and API requests are generated to the generative AI model (e.g., OpenAI's DALL-E or GPT-4).
[2156] Input: Text data received from the device
[2157] Output: API request to the generated AI model
[2158] Step 4:
[2159] Generative AI model: Generates images based on text data. The generated image data is returned to the server.
[2160] Input: Prompt message sent from the server (API request)
[2161] Output: Generated image data
[2162] Step 5:
[2163] Server: Retrieves image data received from the generated AI model and sends it back to the user's terminal. Here, it checks the image data format (e.g., PNG, JPEG) and encodes it in the appropriate format.
[2164] Input: Image data returned from a generative AI model
[2165] Output: Image data to be sent to the user's terminal.
[2166] Step 6:
[2167] Terminal: Renders received image data on a web browser and displays it to the user.
[2168] Input: Image data returned from the server
[2169] Output: Image displayed in the web browser
[2170] Step 7:
[2171] User: Review the generated image and click the "Print" or "Display on TV" button on the web application.
[2172] Input: Image data displayed in a web browser
[2173] Output: User instructions for printing or displaying.
[2174] Step 8:
[2175] Terminal: Triggers user actions (print or view) and sends image data to the server.
[2176] Input: User's print or display instructions
[2177] Output: Image data sent to the server
[2178] Step 9:
[2179] Server: Sends image data to the corresponding device's API to instruct it to print or display. For printing, it sends the data to the printer's API (e.g., HP's Printer API), and for display, it sends the data to the TV's API (e.g., Chromecast API).
[2180] Input: Image data sent from the device
[2181] Output: Image data sent to a printer or television.
[2182] Step 10:
[2183] Printer or television: Receives image data and displays or prints it on the respective device. A printer prints the image on paper, while a television displays the image in high resolution.
[2184] Input: Image data sent from the server
[2185] Output: Printed image or displayed image
[2186] Step 11:
[2187] Server: After the user completes the image generation, printing, and display operations, the server provides a page in the web application that displays product information powered by artificial intelligence technology.
[2188] Input: User operation completion information
[2189] Output: Product information display page
[2190] Step 12:
[2191] User: View information about AI-powered products on a webpage and check the details. Enter questions into the chatbot as needed.
[2192] Input: Viewing product information page
[2193] Output: Question sent to the chatbot
[2194] Step 13:
[2195] Server: Provides automated responses to user questions via a chatbot and transfers the user to a human representative when necessary.
[2196] Input: Question from a user
[2197] Output: Chatbot or human response
[2198] (Application Example 1)
[2199] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2200] Conventional image generation systems had a cumbersome process for users to apply their generated images to specific products and lacked intuitive preview functionality. As a result, users could not visually confirm how the generated images would appear on the actual product, making it difficult to achieve a satisfactory purchasing experience. Furthermore, the means of transmitting the generated images to various output devices for appropriate display and printing were limited.
[2201] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[2202] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for applying the generated images to a specific product, means for displaying the result of applying the generated images to a specific product as a three-dimensional preview, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This allows the user to intuitively apply the generated images to a specific product and confirm them in a three-dimensional preview. Furthermore, the generated images can be easily displayed and printed on various output devices, improving the user's purchasing experience.
[2203] A "means for receiving generation instructions" refers to a device or system that receives input from a user to instruct image generation.
[2204] "Means using a generative artificial intelligence model that generates images based on generation instructions" refers to devices or systems that utilize artificial intelligence technology to generate images based on user instructions.
[2205] "Means for displaying generated images on a user's terminal" refers to devices or systems for displaying generated image data on a user's terminal.
[2206] "Means for transmitting generated image data to a printing device and performing printing" refers to a device or system for transmitting generated image data to a printing device and printing the image using the printing device.
[2207] "Means for transmitting generated image data to a display device and displaying it" refers to a device or system for transmitting generated image data to a display device and displaying the image on the display device.
[2208] "Means for applying a generated image to a specific product" refers to a device or system for applying a generated image to a desired product.
[2209] "Means for displaying the result of applying a generated image to a specific product as a three-dimensional preview" refers to a device or system for applying a generated image to a specific product and visualizing and displaying the result in three dimensions.
[2210] "Means for displaying information about a product equipped with artificial intelligence technology after the generation of the aforementioned image is completed" refers to a device or system for providing information about a product equipped with artificial intelligence technology after the generation process is completed.
[2211] The following describes the specific system configuration and program processing for implementing this invention. In particular, it describes the detailed process of generating images using a generative AI model and applying them to a product.
[2212] Basic System Configuration
[2213] 1. Preparing the user interface
[2214] Server: Starts a web server and hosts a web application for users to access. This web application includes a form for entering instructions for image generation.
[2215] 2. Receive instructions for image generation.
[2216] User: Access the web application and enter the content of the image you want to generate as text. Submit this text instruction by clicking the "Generate" button.
[2217] 3. Image generation using AI
[2218] Server: Analyzes the received text data and sends requests to the generative artificial intelligence model. For example, the text "Colorful geometric patterns" becomes an instruction to generate an image based on its content.
[2219] Server: The generative artificial intelligence model generates images and sends the results back to the server.
[2220] 4. Displaying the generated image
[2221] Server: Sends the received image data back to the user's terminal.
[2222] Terminal: Receives image data and displays it to the user on the web application.
[2223] 5. Application of images to products
[2224] User: Review the generated image and select the option to apply it to a specific product (e.g., T-shirt, cup, etc.).
[2225] Server: Uses image data to place and apply images to specified products.
[2226] Terminal: Displays a 3D preview of the generated product for the user to review.
[2227] 6. Prepare and execute printing.
[2228] User: After reviewing the generated image, decide to order the product and click the "Print" button.
[2229] Server: Receives print instructions and sends the generated image data to the printing device.
[2230] Printing device: Follow the instructions and print the generated image onto the product.
[2231] 7. Preparing and executing the TV display.
[2232] User: After reviewing the generated image, click the "Display on TV" button.
[2233] Server: Triggers a display action and sends image data to the display device.
[2234] Display device: Displays images in high resolution.
[2235] 8. Display of information on AI-equipped products
[2236] Server: After image generation, printing, and display operations are complete, the server displays a page on the web application that provides users with detailed information about the product, which incorporates artificial intelligence technology.
[2237] Hardware and software to be used
[2238] hardware
[2239] User's smartphone
[2240] Server (e.g., cloud service)
[2241] printing device
[2242] Display device (home television)
[2243] software
[2244] Frontend: React
[2245] Backend: Flask (Python)
[2246] Image generation: OpenAI API
[2247] Specific example
[2248] Example of a prompt
[2249] "Colorful geometric patterns"
[2250] "Spacescape"
[2251] "Cute cat illustration"
[2252] For example, a user accesses a web application, enters "colorful geometric patterns," and clicks the "Generate" button. The server, upon receiving this instruction, uses a generative AI model to generate an image based on the specified content. The generated image is first displayed on the user's terminal. Next, the user applies the image to a specific product and checks it in a 3D preview. If the user is satisfied with this preview, they can send the image to a printing device to print it on the product. The generated image can also be displayed on a home television. Finally, after the image generation experience, the user is provided with detailed information about products equipped with artificial intelligence technology.
[2253] In summary, this invention provides a system that allows users to intuitively generate images and quickly apply them to products for visualization. Furthermore, user satisfaction can be enhanced by displaying or printing the generated images on various output devices.
[2254] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[2255] Step 1:
[2256] The user enters a prompt.
[2257] The user accesses the web application and enters a text prompt for image generation (e.g., "Colorful geometric pattern"). Once the input is complete, they click the "Generate" button.
[2258] Input: Text prompt (e.g., "Colorful geometric patterns")
[2259] Output: Generation instructions
[2260] Specific operation: The user enters a prompt message in the text input field and clicks a button, which sends the instruction to the server.
[2261] Step 2:
[2262] The server receives the generation instruction.
[2263] The server receives the prompt message sent by the user and parses it.
[2264] Input: Generation Instructions
[2265] Output: Parsed prompt message
[2266] Specific operation: The server receives an HTTP request and extracts the prompt text.
[2267] Step 3:
[2268] The server sends a request to the generated AI model.
[2269] Based on the parsed prompt text, the server sends an image generation request to the AI model.
[2270] Input: Parsed prompt message
[2271] Output: Request for generated AI model
[2272] Specific operation: The server generates a prompt message and sends it to the AI model's API in the appropriate format.
[2273] Step 4:
[2274] The generative AI model generates images.
[2275] The generative AI model generates an image based on the prompt text and sends the result back to the server.
[2276] Input: Request for generated AI model
[2277] Output: Generated image data
[2278] Specific operation: The generative AI model performs internal processing, generates an image based on the prompt text, and sends that data back to the server.
[2279] Step 5:
[2280] The server receives the generated image and sends it to the user's terminal.
[2281] The server receives the generated image data and sends it back to the user's terminal.
[2282] Input: Generated image data
[2283] Output: Sending image data to the user terminal
[2284] Specific operation: The server sends the received image data to the user's terminal as an HTTP response.
[2285] Step 6:
[2286] Display images on the user's terminal.
[2287] The user terminal displays the received image data on the web application.
[2288] Input: Image data sent to the user terminal
[2289] Output: Image display on a web application
[2290] Specific operation: The user's device browser renders the image data and displays it on the screen.
[2291] Step 7:
[2292] Users apply images to the product.
[2293] The user reviews the generated image and selects the option to apply it to a specific product (e.g., a T-shirt, a cup).
[2294] Input: Generated image and product selection
[2295] Output: Image applied to the product
[2296] Specific operation: The user clicks a product selection option and sends an apply request to the server.
[2297] Step 8:
[2298] The server applies the image to the product.
[2299] The server applies the generated image to a specific product and generates the corresponding data.
[2300] Input: Generated image and product selection information
[2301] Output: Image data applied to the product
[2302] Specific operation: The server places images into the specified product template and generates the applied data.
[2303] Step 9:
[2304] The server sends a 3D preview to the user terminal.
[2305] The server sends the generated product data to the user terminal in a three-dimensional preview format.
[2306] Input: Image data applied to the product
[2307] Output: Sending a 3D preview to the user's terminal
[2308] Specific operation: The server generates data for the 3D preview and sends it to the user's terminal as an HTTP response.
[2309] Step 10:
[2310] The user terminal displays a 3D preview.
[2311] The user terminal displays the received 3D preview on the web application.
[2312] Input: Data for 3D preview
[2313] Output: Three-dimensional preview display on screen
[2314] Specific operation: The user's browser renders a 3D preview and displays it on the screen.
[2315] Step 11:
[2316] The user issues a print command.
[2317] The user reviews the 3D preview, and if satisfied, clicks the "Print" button to initiate printing.
[2318] Input: User's print instructions
[2319] Output: Print request to server
[2320] Specific action: The user clicks a button and sends a print command to the server.
[2321] Step 12:
[2322] The server sends image data to the printer.
[2323] The server receives the print command and sends the generated image data to the printing device.
[2324] Input: User's print instructions and image data.
[2325] Output: Print instructions to the printer
[2326] Specific operation: The server sends image data to the printer's API and issues a print command.
[2327] Step 13:
[2328] The printing device prints the image.
[2329] The printing device prints an image onto the specified product based on the received image data.
[2330] Input: Print instructions to the printer and image data.
[2331] Output: Printed product
[2332] Specific operation: The printing device prints an image onto the product based on the data it receives.
[2333] Step 14:
[2334] The user displays an image on the display device.
[2335] The user clicks the "Display on TV" button to display the generated image on a display device such as a home television.
[2336] Input: User's display instructions
[2337] Output: Display request to the server
[2338] Specific operation: The user clicks a button and sends a display instruction to the server.
[2339] Step 15:
[2340] The server sends image data to the display device.
[2341] The server receives the display instruction and sends the generated image data to the display device.
[2342] Input: User display instructions and image data
[2343] Output: Display instructions for the display device.
[2344] Specific operation: The server sends image data to the display device's API and issues a display instruction.
[2345] Step 16:
[2346] The display device displays an image.
[2347] The display device displays the image in high resolution based on the received image data.
[2348] Input: Display instructions and image data for the display device.
[2349] Output: Displayed image
[2350] Specific operation: The display device displays an image on the screen based on the data it receives.
[2351] Step 17:
[2352] The server displays information about products equipped with AI technology.
[2353] After the server completes image generation, printing, and display operations, it displays a page on the web application that provides users with detailed information about the AI-powered product.
[2354] Input: User's operation complete
[2355] Output: Information on products equipped with AI technology
[2356] Specific action: The server updates the content of the webpage and provides the user with detailed information.
[2357] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[2358] This invention relates to a generative AI system that combines an emotion engine to recognize user emotions. This system receives image generation instructions from the user, generates images using a generative artificial intelligence model, and outputs the generated images to a printer or display device. Furthermore, it displays information about products equipped with artificial intelligence technology after image generation. The system incorporates an emotion engine that recognizes the user's emotional state and adjusts generation instructions based on that state.
[2359] Basic configuration
[2360] 1. Preparing the user interface
[2361] The server starts up the web server and hosts the web application for users to access. The web application displays a form where the user enters instructions for generating an image.
[2362] 2. Initiation of emotion recognition
[2363] The user accesses the web application and clicks a specific button to initiate emotion recognition.
[2364] The device collects user interaction and input data and sends it to the emotion engine.
[2365] 3. Analysis of emotional state
[2366] The server uses an emotion engine to analyze the user's emotional state. For example, it identifies emotions based on the user's input speed, input content, and facial expression data (if acquired using a camera).
[2367] 4. Receive instructions for image generation.
[2368] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[2369] The emotion engine takes the user's emotional state into consideration and makes adaptive adjustments to the generation instructions.
[2370] The terminal sends the adjusted generation instruction data to the server.
[2371] 5. Image generation using AI
[2372] The server analyzes the received text data and sends an image generation request to the generative artificial intelligence model. For example, based on the input text "Sunset Landscape," it generates an image that contains emotionally positive elements.
[2373] 6. Displaying the generated image
[2374] The server receives the generated image data returned from the generated artificial intelligence model and sends it back to the user's terminal.
[2375] The device receives image data and displays it to the user on the web application. The user then views the generated image on the screen.
[2376] 7. Prepare and execute the print job.
[2377] The user reviews the generated image and clicks the "Print" button on the web application.
[2378] The terminal triggers a print action and resends the image data to the server.
[2379] The server receives the image data and sends it to the printer's API to issue a print command.
[2380] The printer prints the image in high quality according to the instructions. The user receives the printed image.
[2381] 8. Preparing and executing the TV display.
[2382] After the user reviews the generated image, they click the "Display on TV" button.
[2383] The device triggers a display action, and the image data is sent to the server again.
[2384] The server sends image data to the TV's API and issues instructions for display.
[2385] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[2386] 9. Information display for AI-equipped products
[2387] After the server has finished generating, printing, and displaying images, it provides users with detailed information about the AI-powered product via a web application.
[2388] The device displays the product information page to the user.
[2389] 10. Considering purchasing a Pixel product
[2390] Users view product information and, if interested, take actions to request additional information.
[2391] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[2392] Specific example
[2393] Example 1: Emotion-driven landscape photography and printing
[2394] The user enters "I want a landscape photo" and clicks the "Generate" button.
[2395] The emotion engine analyzes user input and interaction data to identify positive emotional states.
[2396] The server sends an instruction to the AI model to generate a "positive sunset landscape."
[2397] The server receives the generated landscape photo data and sends it to the terminal.
[2398] The device displays landscape photos it has received.
[2399] After the user reviews the image, they click the "Print" button.
[2400] The server sends the image data to the printer and starts printing.
[2401] The printer prints high-quality landscape photographs.
[2402] Example 2: Generation and display of cat images based on emotions
[2403] The user enters "I want to generate a picture of a cat" and clicks the "Generate" button.
[2404] The emotion engine analyzes the user's interaction data and determines that the user is slightly tired.
[2405] The server sends an instruction to the AI model to generate a "relaxing picture of a cat."
[2406] The server receives the generated cat image data and sends it to the terminal.
[2407] The device displays the image of a cat it received.
[2408] After the user reviews the image, they click the "Display on TV" button.
[2409] The server sends the image data to the television and begins displaying it.
[2410] The TV displays a picture of a cat in high resolution.
[2411] Example 3: Information retrieval for AI-powered products based on emotions
[2412] After the user completes the image generation experience, they can view information pages about AI-powered products on the web application.
[2413] The emotion engine recognizes the user's level of excitement and prioritizes displaying highly relevant product information.
[2414] The server provides detailed product information, and a chatbot answers questions.
[2415] Users can obtain information to consider purchasing and inquire about details as needed.
[2416] In this way, a system is realized that takes user emotions into consideration and provides a consistent service from image generation using generative AI to output methods and product information acquisition. By incorporating an emotion engine, the user experience can be further personalized and satisfaction can be increased.
[2417] The following describes the processing flow.
[2418] Step 1:
[2419] The server starts up a web server and hosts a web application for users to access. This allows users to access a form to enter instructions for image generation.
[2420] Step 2:
[2421] The user accesses a web application and clicks a specific button to initiate emotion recognition. This action causes the emotion engine to begin collecting data about the user.
[2422] Step 3:
[2423] The device collects user interaction data (such as input speed, input content, and in some cases, facial recognition data from the webcam) and sends it to the emotion engine.
[2424] Step 4:
[2425] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the collected data to determine whether the user is in a positive, negative, relaxed, or other emotional state.
[2426] Step 5:
[2427] The user enters the content of the image they want to generate as text and clicks the "Generate" button.
[2428] Step 6:
[2429] The emotion engine considers the user's emotional state and makes adaptive adjustments to the "generate" instructions. For example, if the user is in a positive state, it adds instructions to include vibrant colors and bright scenes in the generated image.
[2430] Step 7:
[2431] The terminal sends the adjusted generation instruction data to the server. The transmitted data includes text as well as adjustment information for the emotion engine.
[2432] Step 8:
[2433] The server analyzes the received text data and sentiment adjustment data, and sends an image generation request to the generative artificial intelligence model. For example, when providing the generative AI model with the text "sunset landscape," instructions are also given to include positive elements.
[2434] Step 9:
[2435] The server receives the generated image data returned from the generated artificial intelligence model and sends it to the user's terminal. The image data is usually encoded in a format such as Base64.
[2436] Step 10:
[2437] The device decodes the received image data and displays it on the web application. The user then views the generated image on the screen.
[2438] Step 11:
[2439] After the user reviews the generated image, they click the "Print" button on the web application.
[2440] Step 12:
[2441] The terminal triggers a print action and resends the image data to the server.
[2442] Step 13:
[2443] The server receives the image data and sends it to the printer's API to issue a print command.
[2444] Step 14:
[2445] The printer follows the instructions and prints the received image data in high quality. The user receives the printed image.
[2446] Step 15:
[2447] After the user reviews the generated image, they click the "Display on TV" button on the web application.
[2448] Step 16:
[2449] The device triggers a display action, and the image data is sent to the server again.
[2450] Step 17:
[2451] The server sends image data to the TV's API and issues instructions for display.
[2452] Step 18:
[2453] The television follows the instructions and displays the received image in high resolution. The user can confirm that the generated image is displayed on a large screen.
[2454] Step 19:
[2455] After the server completes the image generation, printing, and display operations, it provides users with detailed information about the product, which incorporates artificial intelligence technology, through a web application.
[2456] Step 20:
[2457] Users view the displayed product information and, if interested, take action to request additional information.
[2458] Step 21:
[2459] The server provides automated responses via chatbots and handover functions to human staff in response to user actions.
[2460] (Example 2)
[2461] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2462] Conventional image generation systems often fail to consider the user's emotional state when issuing generation instructions, resulting in generated images that do not always match the user's expectations or feelings. Furthermore, the limited output formats and methods for generated images posed a challenge, hindering user convenience.
[2463] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[2464] In this invention, the server includes means for receiving generation instructions, means for using a generation artificial intelligence model that generates images based on the generation instructions, means for displaying the generated images on a user terminal, means for transmitting the generated image data to a printing device and printing, means for transmitting the generated image data to a display device and displaying it, means for adjusting the generation instructions using an emotion engine that recognizes the user's emotional state, and means for displaying information about the product equipped with artificial intelligence technology after the image generation is complete. This makes it possible to generate images that take the user's emotional state into consideration, thereby increasing user satisfaction. In addition, the generated images can be output to a variety of devices, improving user convenience.
[2465] "Generation instructions" are data entered by the user to specify the content and conditions of the image that should be generated.
[2466] A "generative artificial intelligence model" is an artificial intelligence algorithm or system that generates images based on instructions or data input by a user.
[2467] A "user terminal" refers to a device such as a computer or smartphone that a user operates, and which provides an interface for image generation through a web application.
[2468] A "printing device" is a hardware device used to physically print generated images, and primarily refers to a printer.
[2469] A "display device" is a hardware device used to visually present generated images to a user, and includes televisions and monitors.
[2470] An "emotion engine" is software or a system that analyzes user interaction data and input to identify emotional states and adjusts generation instructions based on those results.
[2471] "Information about products equipped with artificial intelligence technology" refers to detailed information about products incorporating artificial intelligence technology, which is provided to the user after the display of the generated image is complete.
[2472] This invention relates to an image generation system that incorporates an emotion engine to recognize user emotions. The system receives image generation instructions from the user and generates images using a generation artificial intelligence model. The generated images are then output to a printer or display device, and information about the product, which incorporates artificial intelligence technology, is displayed after image generation.
[2473] composition
[2474] The system of the present invention has the following components.
[2475] 1. Preparing the user interface
[2476] The server starts an Apache or NGINX web server and hosts a web application for users to input instructions. This application is implemented using HTML / CSS / JavaScript.
[2477] 2. Initiation of emotion recognition
[2478] The user clicks the "Start Emotion Recognition" button in the web application. The device captures the click event and prepares to send interaction data to the emotion engine. This is done using services such as Microsoft Azure's Emotion API.
[2479] 3. Analysis of emotional state
[2480] The server sends the received interaction data to the emotion engine, which then analyzes the user's emotional state. For example, it analyzes input content, input speed, and facial expression data captured by the camera.
[2481] 4. Receive instructions for image generation.
[2482] The user enters the content of the image they want to generate as text and clicks the "Generate" button. The emotion engine considers the emotional state and makes adaptive adjustments to the generation instructions. The device then sends the adjusted generation instruction data to the server.
[2483] ...
Claims
1. A means of receiving generation instructions, A means of using a generative artificial intelligence model that generates images based on generation instructions, A means for displaying the generated image on the user terminal, A means for transmitting the generated image data to a printing device and performing printing, A means for transmitting the generated image data to a display device and displaying it, A system including means for displaying information about a product equipped with artificial intelligence technology after the generation of the aforementioned image is complete.
2. The system according to claim 1, wherein the generating artificial intelligence model generates an image based on text input.
3. The system according to claim 1, comprising means for encoding and decoding image data generated from the artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A