system
The system addresses the limitations of conventional dialogue systems by providing realistic communication and visual expression, allowing users to interact with favorite characters or celebrities through message and image generation, enhancing user engagement.
Patent Information
- Application Number
- JP2024140360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Conventional dialogue systems lack realistic communication and visual expression, limiting user interaction and attachment.
A system that receives user input, analyzes it for intent and emotion, generates responsive messages and images with facial expressions, and sends them to users, enhancing communication with favorite characters or celebrities.
Enables users to enjoy rich, real-time conversations with visual feedback, fostering a sense of attachment and personalization.
Smart Images

Figure 2026037335000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional dialogue systems have difficulty providing realistic communication because their responses to user input messages are limited to standard phrases or limited patterns. Furthermore, they lack image generation capabilities, resulting in a lack of visual expression in dialogue with users. Therefore, there is a demand for dialogue systems that users can enjoy communicating with and develop a sense of attachment to. [Means for solving the problem]
[0005] The present invention provides a system that simultaneously provides realistic communication and visual expression by including a means for receiving an input message from a user and generating a response message based on the content of the message. It also includes a means for generating images with different facial expressions and postures based on the response message, and a means for sending the generated response message and image to the user. This system also includes a function for analyzing the intent and emotion of the message, uploading the image to an external storage device, and sending the URL to the user. This provides a more advanced communication experience that allows users to enjoy conversations with a sense of attachment.
[0006] An "input message" is the text content that a user sends to a dialogue system.
[0007] "Receiving" refers to the terminal or server obtaining an input message.
[0008] "Analysis" refers to understanding the meaning and intent of the content of a received input message using natural language processing technology.
[0009] A "response message" is a reply to the user that is generated based on the analysis results.
[0010] "Generating means" refers to a program or system for automatically generating a response message or image.
[0011] "Facial expressions and posture" refers to the emotions and body postures expressed by the favorite characters or celebrities in the generated images.
[0012] "Image generation AI" refers to artificial intelligence technology that generates images suitable for response messages.
[0013] "Means of sending" refers to the communication technology used to deliver the generated response message or image to the user.
[0014] "External storage device" refers to cloud storage or a database for saving the generated images.
[0015] "URL" refers to an address used to identify and access a particular resource on the Web. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. This system has the function of receiving an input message from the user, generating a response message based on the content of the message, and generating an image that matches the response message and sending it to the user.
[0038] System configuration
[0039] server
[0040] The server has the following main functions:
[0041] 1. Message receiving function: Receives input messages from the user.
[0042] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0043] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0044] 4. Image generation AI function: Generates an image corresponding to the response message.
[0045] 5. Data transmission function: Sends the generated text and image URL to the terminal.
[0046] Terminal
[0047] The terminal performs the following functions:
[0048] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0049] 2. Message sending function: Sends the input message to the server.
[0050] 3. Data reception function: Receives the response message and image URL sent from the server.
[0051] 4. Display function: displays the received text and images to the user.
[0052] user
[0053] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[0054] Operational Overview
[0055] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its contents using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0056] Next, the server uses the dialogue model to generate a response message saying, "I'm feeling great today. How about you?" Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message.
[0057] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[0058] Specific examples
[0059] For example, suppose a user asks, "What's the weather like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. In response, it generates a message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It also generates an image of a person with a cheerful expression that evokes sunny weather. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[0060] In this way, this system allows users to enjoy communicating with their favorite idols both verbally and visually.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] The user enters a message through the app or web interface.
[0064] Step 2:
[0065] The device receives an input message from the user. For example, the user types, "How are you feeling today?"
[0066] Step 3:
[0067] The device sends the received message to the server, including the message and information such as the user ID.
[0068] Step 4:
[0069] The server receives the incoming message.
[0070] Step 5:
[0071] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[0072] Step 6:
[0073] The server uses the dialogue model to generate a response message based on the analysis results, for example, "I'm feeling great today. How about you?"
[0074] Step 7:
[0075] The server calls the image generation AI based on the generated response message.
[0076] Step 8:
[0077] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[0078] Step 9:
[0079] The server uploads the generated images to an external storage device (such as cloud storage).
[0080] Step 10:
[0081] The server retrieves the image URL from cloud storage.
[0082] Step 11:
[0083] The server sends the generated response message and data including the image URL to the terminal.
[0084] Step 12:
[0085] The terminal receives the response message and image URL received from the server.
[0086] Step 13:
[0087] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] Conventional dialogue systems simply respond to messages entered by users with text, making it difficult to provide visually appealing, real-time dialogue. This often results in perfunctory conversations without fully understanding the user's emotions and intentions. Furthermore, conventional systems are limited in the images they generate as responses, making it impossible to generate images that reflect the diverse facial expressions and postures expected by users.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for sending the generated response message and images to the user, means for uploading the images to data storage and sending their URLs to the user, and means for analyzing the message using a natural language processing model and generating images using a generative AI model. This allows the user to receive a response including images with a variety of facial expressions and poses that reflect their emotions and intentions, enabling them to enjoy visually rich, real-time dialogue.
[0093] An "input message" is text data that a user sends through interaction with the system.
[0094] "Analysis" refers to the process performed to understand the meaning and intent of a received message.
[0095] A "response message" is text data that is generated as a response to an input message by the system.
[0096] "Image generation" refers to the process of generating images with specific facial expressions and poses based on text data and analysis results.
[0097] "Data storage" refers to external storage devices and cloud services for saving generated images and data.
[0098] "URL" means a uniform resource identifier for locating resources on the Internet.
[0099] A "natural language processing model" is an algorithm or program that allows a computer to understand and generate human language.
[0100] A "generative AI model" is an artificial intelligence model that generates new data based on input prompts.
[0101] A "server" is a computer system that processes user requests and provides various functions of an interactive system.
[0102] A "terminal" is a device used by a user (e.g., a smartphone or PC) that provides an interface for interacting with the system.
[0103] This invention relates to a system that allows users to enjoy real-time conversations with their favorite characters and celebrities. The system generates appropriate response messages and related images in response to messages entered by users and provides them to the users. This system is primarily composed of a server and a terminal.
[0104] Server configuration and functions
[0105] Message receiving function
[0106] The server has the ability to receive input messages from the user, which can be done using a standard HTTP POST request.
[0107] Message analysis function using natural language processing
[0108] The server analyzes the received message using a natural language processing model (e.g., Hugging Face Transformers) to understand the content and intent of the message.
[0109] Response message generation function
[0110] Based on the analysis results, the server uses a dialogue model (e.g., OpenAI's GPT-3®) to generate an appropriate response message. This process provides an appropriate response to the user's question or comment.
[0111] Image generation function
[0112] Based on the generated response message, an image generation AI (e.g., DALL-E) is used to generate images with different facial expressions and poses, which are then uploaded to cloud storage (e.g., Amazon S3).
[0113] Data transmission function
[0114] The generated response message and image URL are sent to the user's device as a JSON response, allowing the user to receive both the text and the image.
[0115] Device configuration and functions
[0116] Message input function
[0117] The device used by the user (e.g., a smartphone or a computer) provides an interface in which the message can be entered, which includes a text box.
[0118] Message sending function
[0119] The device has the ability to send messages entered by the user to a server, typically using an HTTP POST request.
[0120] Data reception function
[0121] It has the function of receiving response messages and image URLs from the server, and the received data is processed through the application or web browser.
[0122] Display function
[0123] The device displays received text messages and images to the user, allowing the user to enjoy intuitive interaction.
[0124] Specific examples
[0125] For example, here is what happens when a user types "What's the weather like tomorrow?" When a user types "What's the weather like tomorrow?" into a smartphone app, the message is sent to the server via an HTTP POST request. The server analyzes the message using Hugging Face Transformers and recognizes it as a "weather-related question." Next, the server uses OpenAI's GPT-3 to generate a response message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It then uses image generation AI (DALL-E) to generate an image of a "cheerful expression that evokes sunny weather" and uploads it to Amazon S3. Finally, the server sends the response message and the image URL to the user's device, where the user can view it.
[0126] Prompt Sentence Examples
[0127] Here are some example prompts for a generative AI model:
[0128] Natural Language Parsing Prompt:
[0129] User message: "What's the weather like tomorrow?"
[0130] Analysis results: Weather questions
[0131] Dialogue model prompt:
[0132] User message: "What's the weather like tomorrow?"
[0133] Model's response: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0134] Image generation AI prompt:
[0135] Response message: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0136] Generated image: A character image with a cheerful expression that evokes a sunny day
[0137] By combining these functions and processes, the present invention realizes a system that allows users to simultaneously enjoy conversations that reflect their emotions and intentions and visual feedback.
[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0139] Step 1: User message input
[0140] The user inputs a conversation message into the system through a terminal application or a web interface. The input data is a text message, such as "How are you feeling today?" The input message is stored in memory.
[0141] Step 2: Send the message to the server
[0142] The terminal sends the message entered by the user to the server. Specifically, it uses an HTTP POST request to send the entered text message to the server in JSON format. The input data is the user's message, and the output is an HTTP request in JSON format.
[0143] Step 3: Receiving the message on the server
[0144] The server receives messages sent from the device, for example, a request at a specific endpoint using a web framework like Flask. The input data is the request in JSON format, and the output is a text message extracted from that request.
[0145] Step 4: Message analysis using natural language processing
[0146] The server inputs the received message into a natural language processing model (for example, Hugging Face Transformers). The model analyzes the content of the text message and generates data to understand its meaning and intent. The input data is the user's text message, and the output data is the information on intent and meaning as a result of the analysis. Specifically, it analyzes the message "How are you feeling today?" and classifies it as a "question about feelings."
[0147] Step 5: Generate a response message
[0148] The server uses a dialogue model (e.g., OpenAI's GPT-3) based on the analysis results to generate an appropriate response message. The input data is the message analysis result, and the output data is the response message. For example, a response message such as "I'm feeling great today. How about you?" is generated.
[0149] Step 6: Image generation
[0150] Based on the generated response message, the server uses an image generation AI (e.g., DALL-E) to generate a corresponding image. The input data is the response message, and the output data is the generated image. Specifically, an image of a "cheerful expression" is generated and uploaded to cloud storage.
[0151] Step 7: Uploading images to data storage
[0152] The server uploads the generated image to data storage (e.g., Amazon S3) and retrieves its URL. The input data is the generated image, and the output data is the image URL. This allows for long-term storage and access of the image.
[0153] Step 8: Sending response data from the server to the device
[0154] The server sends the generated response message and the image URL together in JSON format to the terminal. The input data is the response message and the image URL, and the output data is the HTTP response in JSON format.
[0155] Step 9: Receive and display data on your device
[0156] The device receives the data sent from the server and displays it to the user. Specifically, it displays the text message "I'm feeling great today. How about you?" and an image of a cheerful expression on the app screen. The input data is the JSON response from the server, and the output data is the text and image displayed to the user.
[0157] In this way, users can enjoy interacting visually and intuitively.
[0158] (Application example 1)
[0159] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0160] Traditional online shopping lacks a way to seamlessly interact with favorite characters or celebrities, a new and unprecedented user experience. This results in a monotonous product browsing experience for users, which reduces the appeal of shopping itself. Furthermore, the lack of real-time customer support and product introductions reduces user choice and satisfaction.
[0161] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0162] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for transmitting the generated response message and image to the user, means for responding in real time to the user's questions and requests so that the user can be guided through the virtual environment, and means for displaying images and information related to the responses. This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and allows them to obtain detailed information about products in real time.
[0163] "Users" refer to people who use the system to shop while enjoying conversations with their favorite characters and celebrities.
[0164] "Input Message" refers to a message such as a question or request that a User sends through the System.
[0165] "Means for receiving" refers to the function of the system to receive an input message sent by a user and send it to a server.
[0166] "Means of analysis" refers to algorithms or functions that understand the received input message and analyze its meaning and intent.
[0167] "Response message" refers to a reply message generated by the system based on the analysis results.
[0168] "Means of generation" refers to the function of using AI or other means to create images with different facial expressions and postures related to the response message.
[0169] "Means for transmitting" refers to a function for transmitting the generated response message and image to provide them to the user.
[0170] "Virtual Environment" refers to the virtual space in which users shop online.
[0171] "Means of responding in real time" refers to the ability to generate a response to user input immediately and engage in dialogue.
[0172] "Means for displaying images and information" refers to the function of visually displaying the generated images and related information to the user.
[0173] MODE FOR CARRYING OUT THE INVENTION
[0174] A specific embodiment of the present invention will now be described. The system of the present invention allows users to enjoy shopping while interacting with their favorite characters or celebrities in a virtual environment. The main configuration and functions of the system are shown below.
[0175] System configuration
[0176] server
[0177] The server has the following main functions:
[0178] 1. Message receiving function: Receives input messages from the user.
[0179] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0180] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0181] 4. Image generation AI function: Generates images with different facial expressions and postures corresponding to the response message.
[0182] 5. Data transmission function: Sends the generated text and image URL to the user's device.
[0183] Terminal
[0184] The terminal performs the following functions:
[0185] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0186] 2. Message sending function: Sends the entered message to the server.
[0187] 3. Data receiving function: Receives the response message and image URL sent from the server.
[0188] 4. Display function: Display the received text and images to the user.
[0189] Processing flow
[0190] The user types a message through the terminal, asking a specific question such as "What are the features of this product?" This generates a prompt like the following:
[0191] Example prompt: What are the features of this product?
[0192] The server receives the message and analyzes it using the content analysis function. Based on the analysis results, the response generation function generates a response message such as "This product has plenty of storage space and a simple, stylish design."
[0193] Next, the image generation AI function generates an image related to the response message (for example, a multi-angle photo of a product) and uploads it to cloud storage. A response message containing the URL of the generated image is sent from the server to the user's device.
[0194] The device receives data from the server and displays an image of the product along with text such as, "This product has ample storage space and a simple, stylish design." This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and obtain detailed information about the product in real time.
[0195] Hardware and software used
[0196] The following hardware and software are used to implement this system:
[0197] Hardware: Smartphone (iOS, ANDROID (registered trademark)), server (cloud-based)
[0198] Software: Python, requests library, Flask (for server-side implementation)
[0199] This allows users to shop in a virtual space while enjoying real-time interactions with their favorite characters and celebrities.
[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0201] Step 1:
[0202] The user inputs a message through the terminal. The input message is a question or request that the user makes to their favorite character or celebrity. For example, a specific question such as "What are the features of this product?" The input message is sent to the server using the message input interface.
[0203] Step 2:
[0204] The server analyzes the received message. It uses content analysis to understand the meaning and intent of the message. A natural language processing (NLP) algorithm is used for this analysis. The message from the user, "What are the features of this product?", is passed as input, and the analysis results are obtained as output. The analysis results include information such as the message being a "question about the product's features."
[0205] Step 3:
[0206] The server generates a response message based on the analysis results. An appropriate reply is generated using the response generation function. For example, based on the analysis results, a response message such as "This product has plenty of storage space and a simple, stylish design" is generated. The input is the analysis results, and the output is the generated response message.
[0207] Step 4:
[0208] The server generates an image based on the generated response message. This process utilizes image generation AI functions. The response message is input as a prompt into the AI model, which then generates an appropriate image (for example, a multi-angle photo of a product). The response message is provided as input, and the generated image is obtained as output.
[0209] Step 5:
[0210] The server uploads the generated image to cloud storage. A URL for the uploaded image is generated and sent to the user device along with a response message. The input is the generated image, and the output is data containing the image URL and the response message.
[0211] Step 6:
[0212] The terminal receives the data sent from the server. Using the data reception function, the response message and image URL are received. The input is the data from the server (the response message and URL), and the output is the received data.
[0213] Step 7:
[0214] The device displays the received response message and image to the user. Using the display function, the text and image are displayed on the user's screen. The received data is the input, and a visual display is the output. This allows the user to enjoy the experience of shopping while interacting with their favorite character or celebrity.
[0215] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0216] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities, and in particular, is capable of recognizing the user's emotions and reflecting them in the content of the interaction and image generation. This system has the function of receiving an input message from the user, generating a response message based on the content and the user's emotions, and further generating an image corresponding to the response message and sending it to the user.
[0217] System configuration
[0218] server
[0219] The server has the following main functions:
[0220] 1. Message receiving function: Receives input messages from the user.
[0221] 2. Content analysis function: Analyzes received messages using natural language processing technology to understand their meaning and intent.
[0222] 3. Emotion recognition function (emotion engine): Analyzes and recognizes the user's emotions from the user's input message.
[0223] 4. Response generation function: Generates an appropriate response message based on the analysis results and the recognized emotions.
[0224] 5. Image generation AI function: Generates images that match the response message and user emotions.
[0225] 6. Data transmission function: Sends the generated text and image URL to the terminal.
[0226] Terminal
[0227] The terminal performs the following functions:
[0228] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0229] 2. Message sending function: Sends the input message to the server.
[0230] 3. Data reception function: Receives the response message and image URL sent from the server.
[0231] 4. Display function: displays the received text and images to the user.
[0232] user
[0233] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[0234] Operational Overview
[0235] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its content using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0236] Next, the server uses an emotion engine to recognize the emotion from the user's input. In this case, the message reads "curiosity." Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[0237] Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and the URL is sent to the device along with the response message.
[0238] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[0239] Specific examples
[0240] For example, if a user asks, "What's the weather going to be like tomorrow?", the server analyzes the message and recognizes that it is a weather-related question. At the same time, the emotion engine recognizes the emotion (e.g., expectation or interest) underlying the user's message.
[0241] In response, a message is generated saying, "It looks like it's going to be sunny tomorrow. I'm looking forward to it!" In addition, an image of a cheerful expression that evokes sunny weather is generated. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[0242] In this way, the system allows users to enjoy both verbal and visual communication with their favorite idols. In particular, the addition of emotion recognition functionality provides more personalized responses, improving the user experience.
[0243] The processing flow will be explained below.
[0244] Step 1:
[0245] The user types a message through the app or web interface, for example, "How are you feeling today?"
[0246] Step 2:
[0247] The terminal receives an input message from the user.
[0248] Step 3:
[0249] The device sends the received message to the server, including the message and information such as the user ID.
[0250] Step 4:
[0251] The server receives the incoming message.
[0252] Step 5:
[0253] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[0254] Step 6:
[0255] The server uses an emotion engine to recognize the user's emotion from the user's input message, for example, recognizing the user's curiosity from the message "How are you feeling today?"
[0256] Step 7:
[0257] The server uses the dialogue model to generate an appropriate response message based on the analysis results and the recognized emotions, for example, "I'm feeling great today. How about you?"
[0258] Step 8:
[0259] The server calls the image generation AI based on the generated response message.
[0260] Step 9:
[0261] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[0262] Step 10:
[0263] The server uploads the generated images to an external storage device (cloud storage).
[0264] Step 11:
[0265] The server retrieves the image URL from cloud storage.
[0266] Step 12:
[0267] The server sends the generated response message and data including the image URL to the terminal.
[0268] Step 13:
[0269] The terminal receives the response message and image URL received from the server.
[0270] Step 14:
[0271] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[0272] Example 2
[0273] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0274] Today's users desire real-time interactive experiences with their favorite characters and celebrities, but conventional dialogue systems lack emotional response and visual feedback. Conventional systems struggle to automatically generate personalized responses and visual expressions that reflect the user's emotions, and improvements to the user experience are needed.
[0275] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message using natural language processing technology and understanding its content, means for recognizing the user's emotion based on the analysis result, means for generating a response message based on the recognized emotion and the analysis result, means for generating an image that matches the emotion based on the response message, and means for sending the URL of the generated response message and image to the user. This allows the user to receive personalized dialogue and visual feedback according to their emotion, improving their real-time dialogue experience.
[0276] A "user" is an individual who uses a system and seeks an interactive experience.
[0277] An "input message" is information in the form of a string that a user sends to the system.
[0278] A "server" is a computer system that receives messages from users, analyzes them, recognizes emotions, generates responses, and generates images.
[0279] "Natural language processing technology" refers to a set of techniques and algorithms that enable computers to understand and analyze human language.
[0280] The "emotion engine" is a software module for recognizing and classifying emotions from user input messages.
[0281] The "response message" is information in the form of a string that the server generates based on the analysis results and the recognized emotion.
[0282] "Image generation AI" is an artificial intelligence technology that generates visual images based on input text and conditions.
[0283] "External storage device" refers to a remote storage service for saving generated data (e.g., image files).
[0284] A "URL" is an address notation for specifying resources on the Internet, and is a type of URI (Uniform Resource Identifier) that indicates the location where the generated image is saved.
[0285] "Means for sending to the user" refers to the communication means for transferring the generated response message and image URL to the user's terminal.
[0286] This invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. Specifically, it has the function of receiving messages input by the user, generating response messages and images based on the content and emotions of the messages, and sending them to the user. This system is mainly composed of three main components: a server, a terminal, and the user.
[0287] System configuration
[0288] server
[0289] The server has the following main functions:
[0290] 1. Message receiving function: Receives input messages from the user. This function can be implemented using, for example, the Python Flask framework.
[0291] 2. Content analysis function: Analyzes received messages using natural language processing technology (e.g., Google (registered trademark) Cloud Natural Language API) to understand their meaning.
[0292] 3. Emotion recognition function (emotion engine): Recognizes emotions based on the analyzed content. This can be implemented using, for example, IBM Watson (registered trademark) Tone Analyzer.
[0293] 4. Response generation function: Generates appropriate response messages based on the recognized emotions and analysis results. This process uses a dialogue generation model such as OpenAI's GPT-3.
[0294] 5. Image generation AI function: Generates images that match the response message and user emotion. For example, DALL·E or similar image generation AI can be used.
[0295] 6. Data transmission function: Sends the generated response message and image URL to the device. Saves the image using a cloud storage service and returns the URL.
[0296] Terminal
[0297] The terminal has the following features:
[0298] 1. Message input function: Providing an interface for users to input messages into a dialogue system, such as a form on a web page or a mobile app.
[0299] 2. Message sending function: Sends the entered message to the server. Data is transferred in JSON format using an HTTP request.
[0300] 3. Data reception function: Receives the response message and image URL sent from the server.
[0301] 4. Display function: Displaying received text and images to the user. This can be achieved on a web page or mobile app using HTML and JavaScript (registered trademark).
[0302] user
[0303] Users interact with the system through an application or web interface. They can type messages and see visual responses and images. For example, if a user types "How are you feeling today?", the server parses the message and generates and sends the appropriate response and associated image.
[0304] Specific examples
[0305] Consider a case where a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. The emotion engine also recognizes the emotion contained in the user's message (for example, anticipation or interest). The response message is generated as "It looks like it's going to be sunny tomorrow. Looking forward to it!" At the same time, an image of a cheerful expression that evokes sunny weather is generated. This response message and the image URL are sent to the user and displayed on the device.
[0306] Prompt Sentence Examples
[0307] Below are some examples of prompts to be used in generative AI models:
[0308] Prompt: "The user asked, 'What's the weather like tomorrow?' They're hoping for sunny weather and would like to see an image of their favorite character with a cheerful expression. Please generate an appropriate response message and image description for this."
[0309] This system allows users to experience emotionally-driven, personalized interactions and visual feedback, enhancing the real-time interaction experience.
[0310] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0311] Step 1:
[0312] The user enters a message into the application or web interface, such as "How are you feeling today?" The input message is stored in text format on the device and sent for further processing.
[0313] Step 2:
[0314] The terminal receives messages entered by the user and sends them to the server through an HTTP request. The messages are converted to JSON format and sent to the server's API endpoint. The input is the user's message, and the output is the JSON data sent to the server.
[0315] Step 3:
[0316] The server receives the message received from the terminal via HTTP request and stores it in the database for analysis. The server is the user's input message and the output is a database storage operation for analysis.
[0317] Step 4:
[0318] The server analyzes the stored messages using natural language processing technology. Specifically, it uses the Google Cloud Natural Language API to understand the meaning and intent of the messages. The input is the user's message read from the database, and the output is the analysis results (e.g., the topic and intent of the message).
[0319] Step 5:
[0320] The server recognizes the user's emotion using an emotion engine (e.g., IBM Watson Tone Analyzer) based on the analysis results. At this stage, the previous analysis results are used as input, and the recognized emotion (e.g., curiosity, anticipation, etc.) is obtained as output.
[0321] Step 6:
[0322] The server uses the analysis results and the recognized emotions as input and generates a response message using a dialogue generation model such as OpenAI GPT-3. The output is the generated response message. For example, a message like "I'm feeling great today. How about you?" is generated.
[0323] Step 7:
[0324] The server generates an image that corresponds to the user's emotion based on the generated response message. It uses image generation AI (e.g., DALL·E) to generate an image that matches the response. The input is the response message and emotion, and the output is the generated image.
[0325] Step 8:
[0326] The server uploads the generated image to an external storage device (e.g., cloud storage) and obtains its URL. The input is the generated image, and the output is the image URL.
[0327] Step 9:
[0328] The server sends the response message and the image URL together in JSON format to the terminal. The input is the response message and the image URL, and the output is the HTTP response sent to the terminal.
[0329] Step 10:
[0330] The terminal analyzes the response message and image URL received from the server and displays them on a web page or application screen. The input is JSON data from the server, and the output is a text message and image displayed to the user.
[0331] (Application example 2)
[0332] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0333] In a dialogue system that allows users to enjoy emotionally rich real-time communication with their favorite characters or celebrities, it is important to accurately recognize the user's emotions and provide response messages and visual feedback that correspond to those emotions. Furthermore, there is a need for a means to efficiently manage the generated images and response messages and provide them to users.
[0334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0335] In this invention, the server includes a means for receiving an input message from a user, a means for analyzing the received message and generating a response message based on the content of the message, a means using artificial intelligence for generating an image based on the response message and the user's emotions, and a means for sending the generated response message and image to the user. This allows users to enjoy emotional communication with their favorite characters or celebrities in real time, and enables personalized responses and visual feedback.
[0336] A "user" is an individual who uses the system of the present invention to enjoy conversations with their favorite characters or celebrities.
[0337] An "input message" is a text-based message that a user sends to the system through a terminal.
[0338] The "receiving means" refers to a function and mechanism for capturing an input message sent by a user into the server.
[0339] The "analysis means" refers to a function and mechanism for analyzing a received message using natural language processing technology and understanding its content and the user's intent.
[0340] A "response message" is a text message generated as a reply by the system based on the content analyzed by the analysis means and the user's recognized emotions.
[0341] The "image generation means" refers to a function and mechanism for generating an appropriate image using artificial intelligence based on the response message and the user's emotions.
[0342] "Artificial intelligence" is a technology that uses large amounts of data to analyze and process user emotions and responses, and then generates images and messages accordingly.
[0343] "Transmission means" refers to the functionality and mechanisms for providing the generated response message and image to the user.
[0344] "Real-time" refers to the timeliness in which a response message and image are generated and sent to the user immediately in response to a user's input.
[0345] The present invention relates to a system that allows users to enjoy emotionally rich communication with their favorite characters and celebrities in real time. Specific embodiments of the system are described below.
[0346] System configuration
[0347] server
[0348] The server has the following main functions:
[0349] Message receiving method:
[0350] Receives input messages from users. At this time, the server receives the user's messages in real time via the Internet.
[0351] Analysis method:
[0352] The received message is analyzed using natural language processing technology, such as software like Google BERT and SpaCy, to understand the meaning and intent of the message.
[0353] Emotion recognition means:
[0354] Analyze and recognize emotions from user input messages. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[0355] Response generation method:
[0356] Based on the analysis results and the recognized emotions, an appropriate response message is generated, using a generative AI model such as GPT-3.
[0357] Image generation method:
[0358] Generate images that match the response message and user emotion using image generation AI such as DALL-E or OpenAI CLIP.
[0359] Data transmission method:
[0360] Send the generated text and image URL to the user's device. Upload the image to cloud storage and obtain its URL.
[0361] Terminal
[0362] The terminal performs the following functions:
[0363] Message input method:
[0364] It provides an interface for users to input messages into the dialogue system.
[0365] Message sending method:
[0366] Sends the input message to the server.
[0367] Data receiving method:
[0368] Receive the response message and image URL sent from the server.
[0369] Display means:
[0370] Display the received text and image to the user.
[0371] Operational Overview
[0372] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, where it is analyzed using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0373] Next, the server uses an emotion recognition means to recognize the emotion from the user's input. In this case, "curiosity" is read from the message. Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[0374] Based on this response message, the image generation means generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message. The device receives the data from the server and displays the text "I'm feeling great today. How about you?" and the image of the favorite character with a cheerful expression to the user.
[0375] Specific examples
[0376] For example, suppose a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. At the same time, the emotion recognition means recognizes the emotion (e.g., anticipation or interest) underlying the user's message. In response, a message such as "It looks like it's going to be sunny tomorrow. Looking forward to it!" is generated. Furthermore, an image of a cheerful expression that evokes the image of sunny weather is generated. This response message and image are sent to the terminal, allowing the user to enjoy it visually.
[0377] This allows the system to allow users to enjoy emotionally rich communication with their favorite characters and celebrities in real time, while providing personalized responses and visual feedback.
[0378] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0379] Step 1:
[0380] The user inputs the message "How are you feeling today?" through the terminal. The terminal receives this input and sends it to the server.
[0381] Input: User input message "How are you feeling today?"
[0382] Output: A request to send a message to the server
[0383] Specific actions: The user enters a message in the device's input interface and presses the send button.
[0384] Step 2:
[0385] The server receives the message sent by the user.
[0386] Input: Message sent from the terminal
[0387] Output: Message data
[0388] Specific operation: The message receiving API on the server receives user messages and stores them in the database.
[0389] Step 3:
[0390] The server parses the received message and performs natural language processing (NLP) to understand its content.
[0391] Input: Received message data
[0392] Output: Analysis results (meaning and intent of the text)
[0393] What it does: Uses an NLP toolkit (e.g. SpaCy, Google BERT) to decompose the message text and analyze its content.
[0394] Step 4:
[0395] The server recognizes the user's emotions from the analysis results and identifies the user's emotions using emotion analysis.
[0396] Input: Analysis results (text meaning and intent)
[0397] Output: Perceived emotion (e.g., curiosity)
[0398] What it does: Analyzes the sentiment of text using a sentiment analysis library (e.g., IBM Watson Tone Analyzer).
[0399] Step 5:
[0400] The server generates a response message based on the analysis results and the recognized emotions.
[0401] Input: Analysis results and recognized emotions
[0402] Output: Response message (e.g. "I'm feeling great today. How about you?")
[0403] Specific operation: Uses response generation AI (e.g., GPT-3) to generate an appropriate response message based on user input.
[0404] Step 6:
[0405] The server uses image generation AI to generate an appropriate image based on the generated response message and the recognized emotion.
[0406] Input: Response message and perceived emotion
[0407] Output: The generated image
[0408] Specific operation: Using image generation AI (e.g., DALL-E, OpenAI CLIP), generate an image corresponding to the response message.
[0409] Step 7:
[0410] The server uploads the generated image to cloud storage and retrieves its URL.
[0411] Input: Generated image
[0412] Output: Image URL
[0413] Specific operation: Calls an API to upload the generated image to cloud storage and obtains the image URL.
[0414] Step 8:
[0415] The server generates a response message and sends the image URL to the terminal.
[0416] Input: Response message and image URL
[0417] Output: Data sent to the terminal
[0418] Specific operation: The server generates a response including a response message and an image URL and sends it to the device.
[0419] Step 9:
[0420] The device receives the data from the server and displays it to the user.
[0421] Input: Response message and image URL sent from the server
[0422] Output: Text and images to display to the user
[0423] Specific operation: The device's receiving interface receives the response from the server and displays text and images on the screen.
[0424] Through these steps, users can enjoy emotionally rich conversations with their favorite characters and celebrities in real time.
[0425] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0426] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0427] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0428] [Second embodiment]
[0429] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0430] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0431] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0432] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0433] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0435] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0436] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0437] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0438] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0439] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0440] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0441] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. This system has the function of receiving an input message from the user, generating a response message based on the content of the message, and generating an image that matches the response message and sending it to the user.
[0442] System configuration
[0443] server
[0444] The server has the following main functions:
[0445] 1. Message receiving function: Receives input messages from the user.
[0446] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0447] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0448] 4. Image generation AI function: Generates an image corresponding to the response message.
[0449] 5. Data transmission function: Sends the generated text and image URL to the terminal.
[0450] Terminal
[0451] The terminal performs the following functions:
[0452] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0453] 2. Message sending function: Sends the input message to the server.
[0454] 3. Data reception function: Receives the response message and image URL sent from the server.
[0455] 4. Display function: displays the received text and images to the user.
[0456] user
[0457] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[0458] Operational Overview
[0459] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its contents using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0460] Next, the server uses the dialogue model to generate a response message saying, "I'm feeling great today. How about you?" Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message.
[0461] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[0462] Specific examples
[0463] For example, suppose a user asks, "What's the weather like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. In response, it generates a message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It also generates an image of a person with a cheerful expression that evokes sunny weather. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[0464] In this way, this system allows users to enjoy communicating with their favorite idols both verbally and visually.
[0465] The processing flow will be explained below.
[0466] Step 1:
[0467] The user enters a message through the app or web interface.
[0468] Step 2:
[0469] The device receives an input message from the user. For example, the user types, "How are you feeling today?"
[0470] Step 3:
[0471] The device sends the received message to the server, including the message and information such as the user ID.
[0472] Step 4:
[0473] The server receives the incoming message.
[0474] Step 5:
[0475] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[0476] Step 6:
[0477] The server uses the dialogue model to generate a response message based on the analysis results, for example, "I'm feeling great today. How about you?"
[0478] Step 7:
[0479] The server calls the image generation AI based on the generated response message.
[0480] Step 8:
[0481] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[0482] Step 9:
[0483] The server uploads the generated images to an external storage device (such as cloud storage).
[0484] Step 10:
[0485] The server retrieves the image URL from cloud storage.
[0486] Step 11:
[0487] The server sends the generated response message and data including the image URL to the terminal.
[0488] Step 12:
[0489] The terminal receives the response message and image URL received from the server.
[0490] Step 13:
[0491] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[0492] Example 1
[0493] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0494] Conventional dialogue systems simply respond to messages entered by users with text, making it difficult to provide visually appealing, real-time dialogue. This often results in perfunctory conversations without fully understanding the user's emotions and intentions. Furthermore, conventional systems are limited in the images they generate as responses, making it impossible to generate images that reflect the diverse facial expressions and postures expected by users.
[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0496] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for sending the generated response message and images to the user, means for uploading the images to data storage and sending their URLs to the user, and means for analyzing the message using a natural language processing model and generating images using a generative AI model. This allows the user to receive a response including images with a variety of facial expressions and poses that reflect their emotions and intentions, enabling them to enjoy visually rich, real-time dialogue.
[0497] An "input message" is text data that a user sends through interaction with the system.
[0498] "Analysis" refers to the process performed to understand the meaning and intent of a received message.
[0499] A "response message" is text data that is generated as a response to an input message by the system.
[0500] "Image generation" refers to the process of generating images with specific facial expressions and poses based on text data and analysis results.
[0501] "Data storage" refers to external storage devices and cloud services for saving generated images and data.
[0502] "URL" means a uniform resource identifier for locating resources on the Internet.
[0503] A "natural language processing model" is an algorithm or program that allows a computer to understand and generate human language.
[0504] A "generative AI model" is an artificial intelligence model that generates new data based on input prompts.
[0505] A "server" is a computer system that processes user requests and provides various functions of an interactive system.
[0506] A "terminal" is a device used by a user (e.g., a smartphone or PC) that provides an interface for interacting with the system.
[0507] This invention relates to a system that allows users to enjoy real-time conversations with their favorite characters and celebrities. The system generates appropriate response messages and related images in response to messages entered by users and provides them to the users. This system is primarily composed of a server and a terminal.
[0508] Server configuration and functions
[0509] Message receiving function
[0510] The server has the ability to receive input messages from the user, which can be done using a standard HTTP POST request.
[0511] Message analysis function using natural language processing
[0512] The server analyzes the received message using a natural language processing model (e.g., Hugging Face Transformers) to understand the content and intent of the message.
[0513] Response message generation function
[0514] Based on the analysis results, the server uses a dialogue model (e.g., OpenAI's GPT-3) to generate an appropriate response message, providing an appropriate response to the user's question or comment.
[0515] Image generation function
[0516] Based on the generated response message, an image generation AI (e.g., DALL-E) is used to generate images with different facial expressions and poses, which are then uploaded to cloud storage (e.g., Amazon S3).
[0517] Data transmission function
[0518] The generated response message and image URL are sent to the user's device as a JSON response, allowing the user to receive both the text and the image.
[0519] Device configuration and functions
[0520] Message input function
[0521] The device used by the user (e.g., a smartphone or a computer) provides an interface in which the message can be entered, which includes a text box.
[0522] Message sending function
[0523] The device has the ability to send messages entered by the user to a server, typically using an HTTP POST request.
[0524] Data reception function
[0525] It has the function of receiving response messages and image URLs from the server, and the received data is processed through the application or web browser.
[0526] Display function
[0527] The device displays received text messages and images to the user, allowing the user to enjoy intuitive interaction.
[0528] Specific examples
[0529] For example, here is what happens when a user types "What's the weather like tomorrow?" When a user types "What's the weather like tomorrow?" into a smartphone app, the message is sent to the server via an HTTP POST request. The server analyzes the message using Hugging Face Transformers and recognizes it as a "weather-related question." Next, the server uses OpenAI's GPT-3 to generate a response message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It then uses image generation AI (DALL-E) to generate an image of a "cheerful expression that evokes sunny weather" and uploads it to Amazon S3. Finally, the server sends the response message and the image URL to the user's device, where the user can view it.
[0530] Prompt Sentence Examples
[0531] Here are some example prompts for a generative AI model:
[0532] Natural Language Parsing Prompt:
[0533] User message: "What's the weather like tomorrow?"
[0534] Analysis results: Weather questions
[0535] Dialogue model prompt:
[0536] User message: "What's the weather like tomorrow?"
[0537] Model's response: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0538] Image generation AI prompt:
[0539] Response message: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0540] Generated image: A character image with a cheerful expression that evokes a sunny day
[0541] By combining these functions and processes, the present invention realizes a system that allows users to simultaneously enjoy conversations that reflect their emotions and intentions and visual feedback.
[0542] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0543] Step 1: User message input
[0544] The user inputs a conversation message into the system through a terminal application or a web interface. The input data is a text message, such as "How are you feeling today?" The input message is stored in memory.
[0545] Step 2: Send the message to the server
[0546] The terminal sends the message entered by the user to the server. Specifically, it uses an HTTP POST request to send the entered text message to the server in JSON format. The input data is the user's message, and the output is an HTTP request in JSON format.
[0547] Step 3: Receiving the message on the server
[0548] The server receives messages sent from the device, for example, a request at a specific endpoint using a web framework like Flask. The input data is the request in JSON format, and the output is a text message extracted from that request.
[0549] Step 4: Message analysis using natural language processing
[0550] The server inputs the received message into a natural language processing model (for example, Hugging Face Transformers). The model analyzes the content of the text message and generates data to understand its meaning and intent. The input data is the user's text message, and the output data is the information on intent and meaning as a result of the analysis. Specifically, it analyzes the message "How are you feeling today?" and classifies it as a "question about feelings."
[0551] Step 5: Generate a response message
[0552] The server uses a dialogue model (e.g., OpenAI's GPT-3) based on the analysis results to generate an appropriate response message. The input data is the message analysis result, and the output data is the response message. For example, a response message such as "I'm feeling great today. How about you?" is generated.
[0553] Step 6: Image generation
[0554] Based on the generated response message, the server uses an image generation AI (e.g., DALL-E) to generate a corresponding image. The input data is the response message, and the output data is the generated image. Specifically, an image of a "cheerful expression" is generated and uploaded to cloud storage.
[0555] Step 7: Uploading images to data storage
[0556] The server uploads the generated image to data storage (e.g., Amazon S3) and retrieves its URL. The input data is the generated image, and the output data is the image URL. This allows for long-term storage and access of the image.
[0557] Step 8: Sending response data from the server to the device
[0558] The server sends the generated response message and the image URL together in JSON format to the terminal. The input data is the response message and the image URL, and the output data is the HTTP response in JSON format.
[0559] Step 9: Receive and display data on your device
[0560] The device receives the data sent from the server and displays it to the user. Specifically, it displays the text message "I'm feeling great today. How about you?" and an image of a cheerful expression on the app screen. The input data is the JSON response from the server, and the output data is the text and image displayed to the user.
[0561] In this way, users can enjoy interacting visually and intuitively.
[0562] (Application example 1)
[0563] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0564] Traditional online shopping lacks a way to seamlessly interact with favorite characters or celebrities, a new and unprecedented user experience. This results in a monotonous product browsing experience for users, which reduces the appeal of shopping itself. Furthermore, the lack of real-time customer support and product introductions reduces user choice and satisfaction.
[0565] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0566] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for transmitting the generated response message and image to the user, means for responding in real time to the user's questions and requests so that the user can be guided through the virtual environment, and means for displaying images and information related to the responses. This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and allows them to obtain detailed information about products in real time.
[0567] "Users" refer to people who use the system to shop while enjoying conversations with their favorite characters and celebrities.
[0568] "Input Message" refers to a message such as a question or request that a User sends through the System.
[0569] "Means for receiving" refers to the function of the system to receive an input message sent by a user and send it to a server.
[0570] "Means of analysis" refers to algorithms or functions that understand the received input message and analyze its meaning and intent.
[0571] "Response message" refers to a reply message generated by the system based on the analysis results.
[0572] "Means of generation" refers to the function of using AI or other means to create images with different facial expressions and postures related to the response message.
[0573] "Means for transmitting" refers to a function for transmitting the generated response message and image to provide them to the user.
[0574] "Virtual Environment" refers to the virtual space in which users shop online.
[0575] "Means of responding in real time" refers to the ability to generate a response to user input immediately and engage in dialogue.
[0576] "Means for displaying images and information" refers to the function of visually displaying the generated images and related information to the user.
[0577] MODE FOR CARRYING OUT THE INVENTION
[0578] A specific embodiment of the present invention will now be described. The system of the present invention allows users to enjoy shopping while interacting with their favorite characters or celebrities in a virtual environment. The main configuration and functions of the system are shown below.
[0579] System configuration
[0580] server
[0581] The server has the following main functions:
[0582] 1. Message receiving function: Receives input messages from the user.
[0583] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0584] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0585] 4. Image generation AI function: Generates images with different facial expressions and postures corresponding to the response message.
[0586] 5. Data transmission function: Sends the generated text and image URL to the user's device.
[0587] Terminal
[0588] The terminal performs the following functions:
[0589] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0590] 2. Message sending function: Sends the entered message to the server.
[0591] 3. Data receiving function: Receives the response message and image URL sent from the server.
[0592] 4. Display function: Display the received text and images to the user.
[0593] Processing flow
[0594] The user types a message through the terminal, asking a specific question such as "What are the features of this product?" This generates a prompt like the following:
[0595] Example prompt: What are the features of this product?
[0596] The server receives the message and analyzes it using the content analysis function. Based on the analysis results, the response generation function generates a response message such as "This product has plenty of storage space and a simple, stylish design."
[0597] Next, the image generation AI function generates an image related to the response message (for example, a multi-angle photo of a product) and uploads it to cloud storage. A response message containing the URL of the generated image is sent from the server to the user's device.
[0598] The device receives data from the server and displays an image of the product along with text such as, "This product has ample storage space and a simple, stylish design." This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and obtain detailed information about the product in real time.
[0599] Hardware and software used
[0600] The following hardware and software are used to implement this system:
[0601] Hardware: Smartphone (iOS, Android), Server (Cloud-based)
[0602] Software: Python, requests library, Flask (for server-side implementation)
[0603] This allows users to shop in a virtual space while enjoying real-time interactions with their favorite characters and celebrities.
[0604] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0605] Step 1:
[0606] The user inputs a message through the terminal. The input message is a question or request that the user makes to their favorite character or celebrity. For example, a specific question such as "What are the features of this product?" The input message is sent to the server using the message input interface.
[0607] Step 2:
[0608] The server analyzes the received message. It uses content analysis to understand the meaning and intent of the message. A natural language processing (NLP) algorithm is used for this analysis. The message from the user, "What are the features of this product?", is passed as input, and the analysis results are obtained as output. The analysis results include information such as the message being a "question about the product's features."
[0609] Step 3:
[0610] The server generates a response message based on the analysis results. An appropriate reply is generated using the response generation function. For example, based on the analysis results, a response message such as "This product has plenty of storage space and a simple, stylish design" is generated. The input is the analysis results, and the output is the generated response message.
[0611] Step 4:
[0612] The server generates an image based on the generated response message. This process utilizes image generation AI functions. The response message is input as a prompt into the AI model, which then generates an appropriate image (for example, a multi-angle photo of a product). The response message is provided as input, and the generated image is obtained as output.
[0613] Step 5:
[0614] The server uploads the generated image to cloud storage. A URL for the uploaded image is generated and sent to the user device along with a response message. The input is the generated image, and the output is data containing the image URL and the response message.
[0615] Step 6:
[0616] The terminal receives the data sent from the server. Using the data reception function, the response message and image URL are received. The input is the data from the server (the response message and URL), and the output is the received data.
[0617] Step 7:
[0618] The device displays the received response message and image to the user. Using the display function, the text and image are displayed on the user's screen. The received data is the input, and a visual display is the output. This allows the user to enjoy the experience of shopping while interacting with their favorite character or celebrity.
[0619] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0620] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities, and in particular, is capable of recognizing the user's emotions and reflecting them in the content of the interaction and image generation. This system has the function of receiving an input message from the user, generating a response message based on the content and the user's emotions, and further generating an image corresponding to the response message and sending it to the user.
[0621] System configuration
[0622] server
[0623] The server has the following main functions:
[0624] 1. Message receiving function: Receives input messages from the user.
[0625] 2. Content analysis function: Analyzes received messages using natural language processing technology to understand their meaning and intent.
[0626] 3. Emotion recognition function (emotion engine): Analyzes and recognizes the user's emotions from the user's input message.
[0627] 4. Response generation function: Generates an appropriate response message based on the analysis results and the recognized emotions.
[0628] 5. Image generation AI function: Generates images that match the response message and user emotions.
[0629] 6. Data transmission function: Sends the generated text and image URL to the terminal.
[0630] Terminal
[0631] The terminal performs the following functions:
[0632] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0633] 2. Message sending function: Sends the input message to the server.
[0634] 3. Data reception function: Receives the response message and image URL sent from the server.
[0635] 4. Display function: displays the received text and images to the user.
[0636] user
[0637] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[0638] Operational Overview
[0639] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its content using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0640] Next, the server uses an emotion engine to recognize the emotion from the user's input. In this case, the message reads "curiosity." Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[0641] Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and the URL is sent to the device along with the response message.
[0642] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[0643] Specific examples
[0644] For example, if a user asks, "What's the weather going to be like tomorrow?", the server analyzes the message and recognizes that it is a weather-related question. At the same time, the emotion engine recognizes the emotion (e.g., expectation or interest) underlying the user's message.
[0645] In response, a message is generated saying, "It looks like it's going to be sunny tomorrow. I'm looking forward to it!" In addition, an image of a cheerful expression that evokes sunny weather is generated. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[0646] In this way, the system allows users to enjoy both verbal and visual communication with their favorite idols. In particular, the addition of emotion recognition functionality provides more personalized responses, improving the user experience.
[0647] The processing flow will be explained below.
[0648] Step 1:
[0649] The user types a message through the app or web interface, for example, "How are you feeling today?"
[0650] Step 2:
[0651] The terminal receives an input message from the user.
[0652] Step 3:
[0653] The device sends the received message to the server, including the message and information such as the user ID.
[0654] Step 4:
[0655] The server receives the incoming message.
[0656] Step 5:
[0657] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[0658] Step 6:
[0659] The server uses an emotion engine to recognize the user's emotion from the user's input message, for example, recognizing the user's curiosity from the message "How are you feeling today?"
[0660] Step 7:
[0661] The server uses the dialogue model to generate an appropriate response message based on the analysis results and the recognized emotions, for example, "I'm feeling great today. How about you?"
[0662] Step 8:
[0663] The server calls the image generation AI based on the generated response message.
[0664] Step 9:
[0665] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[0666] Step 10:
[0667] The server uploads the generated images to an external storage device (cloud storage).
[0668] Step 11:
[0669] The server retrieves the image URL from cloud storage.
[0670] Step 12:
[0671] The server sends the generated response message and data including the image URL to the terminal.
[0672] Step 13:
[0673] The terminal receives the response message and image URL received from the server.
[0674] Step 14:
[0675] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[0676] Example 2
[0677] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0678] Today's users desire real-time interactive experiences with their favorite characters and celebrities, but conventional dialogue systems lack emotional response and visual feedback. Conventional systems struggle to automatically generate personalized responses and visual expressions that reflect the user's emotions, and improvements to the user experience are needed.
[0679] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message using natural language processing technology and understanding its content, means for recognizing the user's emotion based on the analysis result, means for generating a response message based on the recognized emotion and the analysis result, means for generating an image that matches the emotion based on the response message, and means for sending the URL of the generated response message and image to the user. This allows the user to receive personalized dialogue and visual feedback according to their emotion, improving their real-time dialogue experience.
[0680] A "user" is an individual who uses a system and seeks an interactive experience.
[0681] An "input message" is information in the form of a string that a user sends to the system.
[0682] A "server" is a computer system that receives messages from users, analyzes them, recognizes emotions, generates responses, and generates images.
[0683] "Natural language processing technology" refers to a set of techniques and algorithms that enable computers to understand and analyze human language.
[0684] The "emotion engine" is a software module for recognizing and classifying emotions from user input messages.
[0685] The "response message" is information in the form of a string that the server generates based on the analysis results and the recognized emotion.
[0686] "Image generation AI" is an artificial intelligence technology that generates visual images based on input text and conditions.
[0687] "External storage device" refers to a remote storage service for saving generated data (e.g., image files).
[0688] A "URL" is an address notation for specifying resources on the Internet, and is a type of URI (Uniform Resource Identifier) that indicates the location where the generated image is saved.
[0689] "Means for sending to the user" refers to the communication means for transferring the generated response message and image URL to the user's terminal.
[0690] This invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. Specifically, it has the function of receiving messages input by the user, generating response messages and images based on the content and emotions of the messages, and sending them to the user. This system is mainly composed of three main components: a server, a terminal, and the user.
[0691] System configuration
[0692] server
[0693] The server has the following main functions:
[0694] 1. Message receiving function: Receives input messages from the user. This function can be implemented using, for example, the Python Flask framework.
[0695] 2. Content analysis function: Analyzes received messages using natural language processing technology (e.g., Google Cloud Natural Language API) to understand their meaning.
[0696] 3. Emotion recognition function (emotion engine): Recognizes emotions based on the analyzed content. This can be implemented, for example, using IBM Watson Tone Analyzer.
[0697] 4. Response generation function: Generates appropriate response messages based on the recognized emotions and analysis results. This process uses a dialogue generation model such as OpenAI's GPT-3.
[0698] 5. Image generation AI function: Generates images that match the response message and user emotion. For example, DALL·E or similar image generation AI can be used.
[0699] 6. Data transmission function: Sends the generated response message and image URL to the device. Saves the image using a cloud storage service and returns the URL.
[0700] Terminal
[0701] The terminal has the following features:
[0702] 1. Message input function: Providing an interface for users to input messages into a dialogue system, such as a form on a web page or a mobile app.
[0703] 2. Message sending function: Sends the entered message to the server. Data is transferred in JSON format using an HTTP request.
[0704] 3. Data reception function: Receives the response message and image URL sent from the server.
[0705] 4. Display: Displaying the received text and images to the user. This can be achieved using HTML and JavaScript on a web page or mobile app.
[0706] user
[0707] Users interact with the system through an application or web interface. They can type messages and see visual responses and images. For example, if a user types "How are you feeling today?", the server parses the message and generates and sends the appropriate response and associated image.
[0708] Specific examples
[0709] Consider a case where a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. The emotion engine also recognizes the emotion contained in the user's message (for example, anticipation or interest). The response message is generated as "It looks like it's going to be sunny tomorrow. Looking forward to it!" At the same time, an image of a cheerful expression that evokes sunny weather is generated. This response message and the image URL are sent to the user and displayed on the device.
[0710] Prompt Sentence Examples
[0711] Below are some examples of prompts to be used in generative AI models:
[0712] Prompt: "The user asked, 'What's the weather like tomorrow?' They're hoping for sunny weather and would like to see an image of their favorite character with a cheerful expression. Please generate an appropriate response message and image description for this."
[0713] This system allows users to experience emotionally-driven, personalized interactions and visual feedback, enhancing the real-time interaction experience.
[0714] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0715] Step 1:
[0716] The user enters a message into the application or web interface, such as "How are you feeling today?" The input message is stored in text format on the device and sent for further processing.
[0717] Step 2:
[0718] The terminal receives messages entered by the user and sends them to the server through an HTTP request. The messages are converted to JSON format and sent to the server's API endpoint. The input is the user's message, and the output is the JSON data sent to the server.
[0719] Step 3:
[0720] The server receives the message received from the terminal via HTTP request and stores it in the database for analysis. The server is the user's input message and the output is a database storage operation for analysis.
[0721] Step 4:
[0722] The server analyzes the stored messages using natural language processing technology. Specifically, it uses the Google Cloud Natural Language API to understand the meaning and intent of the messages. The input is the user's message read from the database, and the output is the analysis results (e.g., the topic and intent of the message).
[0723] Step 5:
[0724] The server recognizes the user's emotion using an emotion engine (e.g., IBM Watson Tone Analyzer) based on the analysis results. At this stage, the previous analysis results are used as input, and the recognized emotion (e.g., curiosity, anticipation, etc.) is obtained as output.
[0725] Step 6:
[0726] The server uses the analysis results and the recognized emotions as input and generates a response message using a dialogue generation model such as OpenAI GPT-3. The output is the generated response message. For example, a message like "I'm feeling great today. How about you?" is generated.
[0727] Step 7:
[0728] The server generates an image that corresponds to the user's emotion based on the generated response message. It uses image generation AI (e.g., DALL·E) to generate an image that matches the response. The input is the response message and emotion, and the output is the generated image.
[0729] Step 8:
[0730] The server uploads the generated image to an external storage device (e.g., cloud storage) and obtains its URL. The input is the generated image, and the output is the image URL.
[0731] Step 9:
[0732] The server sends the response message and the image URL together in JSON format to the terminal. The input is the response message and the image URL, and the output is the HTTP response sent to the terminal.
[0733] Step 10:
[0734] The terminal analyzes the response message and image URL received from the server and displays them on a web page or application screen. The input is JSON data from the server, and the output is a text message and image displayed to the user.
[0735] (Application example 2)
[0736] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0737] In a dialogue system that allows users to enjoy emotionally rich real-time communication with their favorite characters or celebrities, it is important to accurately recognize the user's emotions and provide response messages and visual feedback that correspond to those emotions. Furthermore, there is a need for a means to efficiently manage the generated images and response messages and provide them to users.
[0738] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0739] In this invention, the server includes a means for receiving an input message from a user, a means for analyzing the received message and generating a response message based on the content of the message, a means using artificial intelligence for generating an image based on the response message and the user's emotions, and a means for sending the generated response message and image to the user. This allows users to enjoy emotional communication with their favorite characters or celebrities in real time, and enables personalized responses and visual feedback.
[0740] A "user" is an individual who uses the system of the present invention to enjoy conversations with their favorite characters or celebrities.
[0741] An "input message" is a text-based message that a user sends to the system through a terminal.
[0742] The "receiving means" refers to a function and mechanism for capturing an input message sent by a user into the server.
[0743] The "analysis means" refers to a function and mechanism for analyzing a received message using natural language processing technology and understanding its content and the user's intent.
[0744] A "response message" is a text message generated as a reply by the system based on the content analyzed by the analysis means and the user's recognized emotions.
[0745] The "image generation means" refers to a function and mechanism for generating an appropriate image using artificial intelligence based on the response message and the user's emotions.
[0746] "Artificial intelligence" is a technology that uses large amounts of data to analyze and process user emotions and responses, and then generates images and messages accordingly.
[0747] "Transmission means" refers to the functionality and mechanisms for providing the generated response message and image to the user.
[0748] "Real-time" refers to the timeliness in which a response message and image are generated and sent to the user immediately in response to a user's input.
[0749] The present invention relates to a system that allows users to enjoy emotionally rich communication with their favorite characters and celebrities in real time. Specific embodiments of the system are described below.
[0750] System configuration
[0751] server
[0752] The server has the following main functions:
[0753] Message receiving method:
[0754] Receives input messages from users. At this time, the server receives the user's messages in real time via the Internet.
[0755] Analysis method:
[0756] The received message is analyzed using natural language processing technology, such as software like Google BERT and SpaCy, to understand the meaning and intent of the message.
[0757] Emotion recognition means:
[0758] Analyze and recognize emotions from user input messages. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[0759] Response generation method:
[0760] Based on the analysis results and the recognized emotions, an appropriate response message is generated, using a generative AI model such as GPT-3.
[0761] Image generation method:
[0762] Generate images that match the response message and user emotion using image generation AI such as DALL-E or OpenAI CLIP.
[0763] Data transmission method:
[0764] Send the generated text and image URL to the user's device. Upload the image to cloud storage and obtain its URL.
[0765] Terminal
[0766] The terminal performs the following functions:
[0767] Message input method:
[0768] It provides an interface for users to input messages into the dialogue system.
[0769] Message sending method:
[0770] Sends the input message to the server.
[0771] Data receiving method:
[0772] Receive the response message and image URL sent from the server.
[0773] Display means:
[0774] Display the received text and image to the user.
[0775] Operational Overview
[0776] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, where it is analyzed using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0777] Next, the server uses an emotion recognition means to recognize the emotion from the user's input. In this case, "curiosity" is read from the message. Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[0778] Based on this response message, the image generation means generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message. The device receives the data from the server and displays the text "I'm feeling great today. How about you?" and the image of the favorite character with a cheerful expression to the user.
[0779] Specific examples
[0780] For example, suppose a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. At the same time, the emotion recognition means recognizes the emotion (e.g., anticipation or interest) underlying the user's message. In response, a message such as "It looks like it's going to be sunny tomorrow. Looking forward to it!" is generated. Furthermore, an image of a cheerful expression that evokes the image of sunny weather is generated. This response message and image are sent to the terminal, allowing the user to enjoy it visually.
[0781] This allows the system to allow users to enjoy emotionally rich communication with their favorite characters and celebrities in real time, while providing personalized responses and visual feedback.
[0782] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0783] Step 1:
[0784] The user inputs the message "How are you feeling today?" through the terminal. The terminal receives this input and sends it to the server.
[0785] Input: User input message "How are you feeling today?"
[0786] Output: A request to send a message to the server
[0787] Specific actions: The user enters a message in the device's input interface and presses the send button.
[0788] Step 2:
[0789] The server receives the message sent by the user.
[0790] Input: Message sent from the terminal
[0791] Output: Message data
[0792] Specific operation: The message receiving API on the server receives user messages and stores them in the database.
[0793] Step 3:
[0794] The server parses the received message and performs natural language processing (NLP) to understand its content.
[0795] Input: Received message data
[0796] Output: Analysis results (meaning and intent of the text)
[0797] What it does: Uses an NLP toolkit (e.g. SpaCy, Google BERT) to decompose the message text and analyze its content.
[0798] Step 4:
[0799] The server recognizes the user's emotions from the analysis results and identifies the user's emotions using emotion analysis.
[0800] Input: Analysis results (text meaning and intent)
[0801] Output: Perceived emotion (e.g., curiosity)
[0802] What it does: Analyzes the sentiment of text using a sentiment analysis library (e.g., IBM Watson Tone Analyzer).
[0803] Step 5:
[0804] The server generates a response message based on the analysis results and the recognized emotions.
[0805] Input: Analysis results and recognized emotions
[0806] Output: Response message (e.g. "I'm feeling great today. How about you?")
[0807] Specific operation: Uses response generation AI (e.g., GPT-3) to generate an appropriate response message based on user input.
[0808] Step 6:
[0809] The server uses image generation AI to generate an appropriate image based on the generated response message and the recognized emotion.
[0810] Input: Response message and perceived emotion
[0811] Output: The generated image
[0812] Specific operation: Using image generation AI (e.g., DALL-E, OpenAI CLIP), generate an image corresponding to the response message.
[0813] Step 7:
[0814] The server uploads the generated image to cloud storage and retrieves its URL.
[0815] Input: Generated image
[0816] Output: Image URL
[0817] Specific operation: Calls an API to upload the generated image to cloud storage and obtains the image URL.
[0818] Step 8:
[0819] The server generates a response message and sends the image URL to the terminal.
[0820] Input: Response message and image URL
[0821] Output: Data sent to the terminal
[0822] Specific operation: The server generates a response including a response message and an image URL and sends it to the device.
[0823] Step 9:
[0824] The device receives the data from the server and displays it to the user.
[0825] Input: Response message and image URL sent from the server
[0826] Output: Text and images to display to the user
[0827] Specific operation: The device's receiving interface receives the response from the server and displays text and images on the screen.
[0828] Through these steps, users can enjoy emotionally rich conversations with their favorite characters and celebrities in real time.
[0829] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0830] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0831] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0832] [Third embodiment]
[0833] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0834] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0835] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0836] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0837] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0838] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0839] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0840] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0841] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0842] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0843] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0844] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0845] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. This system has the function of receiving an input message from the user, generating a response message based on the content of the message, and generating an image that matches the response message and sending it to the user.
[0846] System configuration
[0847] server
[0848] The server has the following main functions:
[0849] 1. Message receiving function: Receives input messages from the user.
[0850] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0851] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0852] 4. Image generation AI function: Generates an image corresponding to the response message.
[0853] 5. Data transmission function: Sends the generated text and image URL to the terminal.
[0854] Terminal
[0855] The terminal performs the following functions:
[0856] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0857] 2. Message sending function: Sends the input message to the server.
[0858] 3. Data reception function: Receives the response message and image URL sent from the server.
[0859] 4. Display function: displays the received text and images to the user.
[0860] user
[0861] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[0862] Operational Overview
[0863] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its contents using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[0864] Next, the server uses the dialogue model to generate a response message saying, "I'm feeling great today. How about you?" Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message.
[0865] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[0866] Specific examples
[0867] For example, suppose a user asks, "What's the weather like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. In response, it generates a message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It also generates an image of a person with a cheerful expression that evokes sunny weather. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[0868] In this way, this system allows users to enjoy communicating with their favorite idols both verbally and visually.
[0869] The processing flow will be explained below.
[0870] Step 1:
[0871] The user enters a message through the app or web interface.
[0872] Step 2:
[0873] The device receives an input message from the user. For example, the user types, "How are you feeling today?"
[0874] Step 3:
[0875] The device sends the received message to the server, including the message and information such as the user ID.
[0876] Step 4:
[0877] The server receives the incoming message.
[0878] Step 5:
[0879] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[0880] Step 6:
[0881] The server uses the dialogue model to generate a response message based on the analysis results, for example, "I'm feeling great today. How about you?"
[0882] Step 7:
[0883] The server calls the image generation AI based on the generated response message.
[0884] Step 8:
[0885] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[0886] Step 9:
[0887] The server uploads the generated images to an external storage device (such as cloud storage).
[0888] Step 10:
[0889] The server retrieves the image URL from cloud storage.
[0890] Step 11:
[0891] The server sends the generated response message and data including the image URL to the terminal.
[0892] Step 12:
[0893] The terminal receives the response message and image URL received from the server.
[0894] Step 13:
[0895] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[0896] Example 1
[0897] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0898] Conventional dialogue systems simply respond to messages entered by users with text, making it difficult to provide visually appealing, real-time dialogue. This often results in perfunctory conversations without fully understanding the user's emotions and intentions. Furthermore, conventional systems are limited in the images they generate as responses, making it impossible to generate images that reflect the diverse facial expressions and postures expected by users.
[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0900] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for sending the generated response message and images to the user, means for uploading the images to data storage and sending their URLs to the user, and means for analyzing the message using a natural language processing model and generating images using a generative AI model. This allows the user to receive a response including images with a variety of facial expressions and poses that reflect their emotions and intentions, enabling them to enjoy visually rich, real-time dialogue.
[0901] An "input message" is text data that a user sends through interaction with the system.
[0902] "Analysis" refers to the process performed to understand the meaning and intent of a received message.
[0903] A "response message" is text data that is generated as a response to an input message by the system.
[0904] "Image generation" refers to the process of generating images with specific facial expressions and poses based on text data and analysis results.
[0905] "Data storage" refers to external storage devices and cloud services for saving generated images and data.
[0906] "URL" means a uniform resource identifier for locating resources on the Internet.
[0907] A "natural language processing model" is an algorithm or program that allows a computer to understand and generate human language.
[0908] A "generative AI model" is an artificial intelligence model that generates new data based on input prompts.
[0909] A "server" is a computer system that processes user requests and provides various functions of an interactive system.
[0910] A "terminal" is a device used by a user (e.g., a smartphone or PC) that provides an interface for interacting with the system.
[0911] This invention relates to a system that allows users to enjoy real-time conversations with their favorite characters and celebrities. The system generates appropriate response messages and related images in response to messages entered by users and provides them to the users. This system is primarily composed of a server and a terminal.
[0912] Server configuration and functions
[0913] Message receiving function
[0914] The server has the ability to receive input messages from the user, which can be done using a standard HTTP POST request.
[0915] Message analysis function using natural language processing
[0916] The server analyzes the received message using a natural language processing model (e.g., Hugging Face Transformers) to understand the content and intent of the message.
[0917] Response message generation function
[0918] Based on the analysis results, the server uses a dialogue model (e.g., OpenAI's GPT-3) to generate an appropriate response message, providing an appropriate response to the user's question or comment.
[0919] Image generation function
[0920] Based on the generated response message, an image generation AI (e.g., DALL-E) is used to generate images with different facial expressions and poses, which are then uploaded to cloud storage (e.g., Amazon S3).
[0921] Data transmission function
[0922] The generated response message and image URL are sent to the user's device as a JSON response, allowing the user to receive both the text and the image.
[0923] Device configuration and functions
[0924] Message input function
[0925] The device used by the user (e.g., a smartphone or a computer) provides an interface in which the message can be entered, which includes a text box.
[0926] Message sending function
[0927] The device has the ability to send messages entered by the user to a server, typically using an HTTP POST request.
[0928] Data reception function
[0929] It has the function of receiving response messages and image URLs from the server, and the received data is processed through the application or web browser.
[0930] Display function
[0931] The device displays received text messages and images to the user, allowing the user to enjoy intuitive interaction.
[0932] Specific examples
[0933] For example, here is what happens when a user types "What's the weather like tomorrow?" When a user types "What's the weather like tomorrow?" into a smartphone app, the message is sent to the server via an HTTP POST request. The server analyzes the message using Hugging Face Transformers and recognizes it as a "weather-related question." Next, the server uses OpenAI's GPT-3 to generate a response message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It then uses image generation AI (DALL-E) to generate an image of a "cheerful expression that evokes sunny weather" and uploads it to Amazon S3. Finally, the server sends the response message and the image URL to the user's device, where the user can view it.
[0934] Prompt Sentence Examples
[0935] Here are some example prompts for a generative AI model:
[0936] Natural Language Parsing Prompt:
[0937] User message: "What's the weather like tomorrow?"
[0938] Analysis results: Weather questions
[0939] Dialogue model prompt:
[0940] User message: "What's the weather like tomorrow?"
[0941] Model's response: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0942] Image generation AI prompt:
[0943] Response message: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[0944] Generated image: A character image with a cheerful expression that evokes a sunny day
[0945] By combining these functions and processes, the present invention realizes a system that allows users to simultaneously enjoy conversations that reflect their emotions and intentions and visual feedback.
[0946] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0947] Step 1: User message input
[0948] The user inputs a conversation message into the system through a terminal application or a web interface. The input data is a text message, such as "How are you feeling today?" The input message is stored in memory.
[0949] Step 2: Send the message to the server
[0950] The terminal sends the message entered by the user to the server. Specifically, it uses an HTTP POST request to send the entered text message to the server in JSON format. The input data is the user's message, and the output is an HTTP request in JSON format.
[0951] Step 3: Receiving the message on the server
[0952] The server receives messages sent from the device, for example, a request at a specific endpoint using a web framework like Flask. The input data is the request in JSON format, and the output is a text message extracted from that request.
[0953] Step 4: Message analysis using natural language processing
[0954] The server inputs the received message into a natural language processing model (for example, Hugging Face Transformers). The model analyzes the content of the text message and generates data to understand its meaning and intent. The input data is the user's text message, and the output data is the information on intent and meaning as a result of the analysis. Specifically, it analyzes the message "How are you feeling today?" and classifies it as a "question about feelings."
[0955] Step 5: Generate a response message
[0956] The server uses a dialogue model (e.g., OpenAI's GPT-3) based on the analysis results to generate an appropriate response message. The input data is the message analysis result, and the output data is the response message. For example, a response message such as "I'm feeling great today. How about you?" is generated.
[0957] Step 6: Image generation
[0958] Based on the generated response message, the server uses an image generation AI (e.g., DALL-E) to generate a corresponding image. The input data is the response message, and the output data is the generated image. Specifically, an image of a "cheerful expression" is generated and uploaded to cloud storage.
[0959] Step 7: Uploading images to data storage
[0960] The server uploads the generated image to data storage (e.g., Amazon S3) and retrieves its URL. The input data is the generated image, and the output data is the image URL. This allows for long-term storage and access of the image.
[0961] Step 8: Sending response data from the server to the device
[0962] The server sends the generated response message and the image URL together in JSON format to the terminal. The input data is the response message and the image URL, and the output data is the HTTP response in JSON format.
[0963] Step 9: Receive and display data on your device
[0964] The device receives the data sent from the server and displays it to the user. Specifically, it displays the text message "I'm feeling great today. How about you?" and an image of a cheerful expression on the app screen. The input data is the JSON response from the server, and the output data is the text and image displayed to the user.
[0965] In this way, users can enjoy interacting visually and intuitively.
[0966] (Application example 1)
[0967] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0968] Traditional online shopping lacks a way to seamlessly interact with favorite characters or celebrities, a new and unprecedented user experience. This results in a monotonous product browsing experience for users, which reduces the appeal of shopping itself. Furthermore, the lack of real-time customer support and product introductions reduces user choice and satisfaction.
[0969] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0970] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for transmitting the generated response message and image to the user, means for responding in real time to the user's questions and requests so that the user can be guided through the virtual environment, and means for displaying images and information related to the responses. This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and allows them to obtain detailed information about products in real time.
[0971] "Users" refer to people who use the system to shop while enjoying conversations with their favorite characters and celebrities.
[0972] "Input Message" refers to a message such as a question or request that a User sends through the System.
[0973] "Means for receiving" refers to the function of the system to receive an input message sent by a user and send it to a server.
[0974] "Means of analysis" refers to algorithms or functions that understand the received input message and analyze its meaning and intent.
[0975] "Response message" refers to a reply message generated by the system based on the analysis results.
[0976] "Means of generation" refers to the function of using AI or other means to create images with different facial expressions and postures related to the response message.
[0977] "Means for transmitting" refers to a function for transmitting the generated response message and image to provide them to the user.
[0978] "Virtual Environment" refers to the virtual space in which users shop online.
[0979] "Means of responding in real time" refers to the ability to generate a response to user input immediately and engage in dialogue.
[0980] "Means for displaying images and information" refers to the function of visually displaying the generated images and related information to the user.
[0981] MODE FOR CARRYING OUT THE INVENTION
[0982] A specific embodiment of the present invention will now be described. The system of the present invention allows users to enjoy shopping while interacting with their favorite characters or celebrities in a virtual environment. The main configuration and functions of the system are shown below.
[0983] System configuration
[0984] server
[0985] The server has the following main functions:
[0986] 1. Message receiving function: Receives input messages from the user.
[0987] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[0988] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[0989] 4. Image generation AI function: Generates images with different facial expressions and postures corresponding to the response message.
[0990] 5. Data transmission function: Sends the generated text and image URL to the user's device.
[0991] Terminal
[0992] The terminal performs the following functions:
[0993] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[0994] 2. Message sending function: Sends the entered message to the server.
[0995] 3. Data receiving function: Receives the response message and image URL sent from the server.
[0996] 4. Display function: Display the received text and images to the user.
[0997] Processing flow
[0998] The user types a message through the terminal, asking a specific question such as "What are the features of this product?" This generates a prompt like the following:
[0999] Example prompt: What are the features of this product?
[1000] The server receives the message and analyzes it using the content analysis function. Based on the analysis results, the response generation function generates a response message such as "This product has plenty of storage space and a simple, stylish design."
[1001] Next, the image generation AI function generates an image related to the response message (for example, a multi-angle photo of a product) and uploads it to cloud storage. A response message containing the URL of the generated image is sent from the server to the user's device.
[1002] The device receives data from the server and displays an image of the product along with text such as, "This product has ample storage space and a simple, stylish design." This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and obtain detailed information about the product in real time.
[1003] Hardware and software used
[1004] The following hardware and software are used to implement this system:
[1005] Hardware: Smartphone (iOS, Android), Server (Cloud-based)
[1006] Software: Python, requests library, Flask (for server-side implementation)
[1007] This allows users to shop in a virtual space while enjoying real-time interactions with their favorite characters and celebrities.
[1008] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1009] Step 1:
[1010] The user inputs a message through the terminal. The input message is a question or request that the user makes to their favorite character or celebrity. For example, a specific question such as "What are the features of this product?" The input message is sent to the server using the message input interface.
[1011] Step 2:
[1012] The server analyzes the received message. It uses content analysis to understand the meaning and intent of the message. A natural language processing (NLP) algorithm is used for this analysis. The message from the user, "What are the features of this product?", is passed as input, and the analysis results are obtained as output. The analysis results include information such as the message being a "question about the product's features."
[1013] Step 3:
[1014] The server generates a response message based on the analysis results. An appropriate reply is generated using the response generation function. For example, based on the analysis results, a response message such as "This product has plenty of storage space and a simple, stylish design" is generated. The input is the analysis results, and the output is the generated response message.
[1015] Step 4:
[1016] The server generates an image based on the generated response message. This process utilizes image generation AI functions. The response message is input as a prompt into the AI model, which then generates an appropriate image (for example, a multi-angle photo of a product). The response message is provided as input, and the generated image is obtained as output.
[1017] Step 5:
[1018] The server uploads the generated image to cloud storage. A URL for the uploaded image is generated and sent to the user device along with a response message. The input is the generated image, and the output is data containing the image URL and the response message.
[1019] Step 6:
[1020] The terminal receives the data sent from the server. Using the data reception function, the response message and image URL are received. The input is the data from the server (the response message and URL), and the output is the received data.
[1021] Step 7:
[1022] The device displays the received response message and image to the user. Using the display function, the text and image are displayed on the user's screen. The received data is the input, and a visual display is the output. This allows the user to enjoy the experience of shopping while interacting with their favorite character or celebrity.
[1023] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1024] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities, and in particular, is capable of recognizing the user's emotions and reflecting them in the content of the interaction and image generation. This system has the function of receiving an input message from the user, generating a response message based on the content and the user's emotions, and further generating an image corresponding to the response message and sending it to the user.
[1025] System configuration
[1026] server
[1027] The server has the following main functions:
[1028] 1. Message receiving function: Receives input messages from the user.
[1029] 2. Content analysis function: Analyzes received messages using natural language processing technology to understand their meaning and intent.
[1030] 3. Emotion recognition function (emotion engine): Analyzes and recognizes the user's emotions from the user's input message.
[1031] 4. Response generation function: Generates an appropriate response message based on the analysis results and the recognized emotions.
[1032] 5. Image generation AI function: Generates images that match the response message and user emotions.
[1033] 6. Data transmission function: Sends the generated text and image URL to the terminal.
[1034] Terminal
[1035] The terminal performs the following functions:
[1036] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[1037] 2. Message sending function: Sends the input message to the server.
[1038] 3. Data reception function: Receives the response message and image URL sent from the server.
[1039] 4. Display function: displays the received text and images to the user.
[1040] user
[1041] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[1042] Operational Overview
[1043] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its content using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[1044] Next, the server uses an emotion engine to recognize the emotion from the user's input. In this case, the message reads "curiosity." Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[1045] Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and the URL is sent to the device along with the response message.
[1046] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[1047] Specific examples
[1048] For example, if a user asks, "What's the weather going to be like tomorrow?", the server analyzes the message and recognizes that it is a weather-related question. At the same time, the emotion engine recognizes the emotion (e.g., expectation or interest) underlying the user's message.
[1049] In response, a message is generated saying, "It looks like it's going to be sunny tomorrow. I'm looking forward to it!" In addition, an image of a cheerful expression that evokes sunny weather is generated. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[1050] In this way, the system allows users to enjoy both verbal and visual communication with their favorite idols. In particular, the addition of emotion recognition functionality provides more personalized responses, improving the user experience.
[1051] The processing flow will be explained below.
[1052] Step 1:
[1053] The user types a message through the app or web interface, for example, "How are you feeling today?"
[1054] Step 2:
[1055] The terminal receives an input message from the user.
[1056] Step 3:
[1057] The device sends the received message to the server, including the message and information such as the user ID.
[1058] Step 4:
[1059] The server receives the incoming message.
[1060] Step 5:
[1061] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[1062] Step 6:
[1063] The server uses an emotion engine to recognize the user's emotion from the user's input message, for example, recognizing the user's curiosity from the message "How are you feeling today?"
[1064] Step 7:
[1065] The server uses the dialogue model to generate an appropriate response message based on the analysis results and the recognized emotions, for example, "I'm feeling great today. How about you?"
[1066] Step 8:
[1067] The server calls the image generation AI based on the generated response message.
[1068] Step 9:
[1069] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[1070] Step 10:
[1071] The server uploads the generated images to an external storage device (cloud storage).
[1072] Step 11:
[1073] The server retrieves the image URL from cloud storage.
[1074] Step 12:
[1075] The server sends the generated response message and data including the image URL to the terminal.
[1076] Step 13:
[1077] The terminal receives the response message and image URL received from the server.
[1078] Step 14:
[1079] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[1080] Example 2
[1081] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1082] Today's users desire real-time interactive experiences with their favorite characters and celebrities, but conventional dialogue systems lack emotional response and visual feedback. Conventional systems struggle to automatically generate personalized responses and visual expressions that reflect the user's emotions, and improvements to the user experience are needed.
[1083] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message using natural language processing technology and understanding its content, means for recognizing the user's emotion based on the analysis result, means for generating a response message based on the recognized emotion and the analysis result, means for generating an image that matches the emotion based on the response message, and means for sending the URL of the generated response message and image to the user. This allows the user to receive personalized dialogue and visual feedback according to their emotion, improving their real-time dialogue experience.
[1084] A "user" is an individual who uses a system and seeks an interactive experience.
[1085] An "input message" is information in the form of a string that a user sends to the system.
[1086] A "server" is a computer system that receives messages from users, analyzes them, recognizes emotions, generates responses, and generates images.
[1087] "Natural language processing technology" refers to a set of techniques and algorithms that enable computers to understand and analyze human language.
[1088] The "emotion engine" is a software module for recognizing and classifying emotions from user input messages.
[1089] The "response message" is information in the form of a string that the server generates based on the analysis results and the recognized emotion.
[1090] "Image generation AI" is an artificial intelligence technology that generates visual images based on input text and conditions.
[1091] "External storage device" refers to a remote storage service for saving generated data (e.g., image files).
[1092] A "URL" is an address notation for specifying resources on the Internet, and is a type of URI (Uniform Resource Identifier) that indicates the location where the generated image is saved.
[1093] "Means for sending to the user" refers to the communication means for transferring the generated response message and image URL to the user's terminal.
[1094] This invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. Specifically, it has the function of receiving messages input by the user, generating response messages and images based on the content and emotions of the messages, and sending them to the user. This system is mainly composed of three main components: a server, a terminal, and the user.
[1095] System configuration
[1096] server
[1097] The server has the following main functions:
[1098] 1. Message receiving function: Receives input messages from the user. This function can be implemented using, for example, the Python Flask framework.
[1099] 2. Content analysis function: Analyzes received messages using natural language processing technology (e.g., Google Cloud Natural Language API) to understand their meaning.
[1100] 3. Emotion recognition function (emotion engine): Recognizes emotions based on the analyzed content. This can be implemented, for example, using IBM Watson Tone Analyzer.
[1101] 4. Response generation function: Generates appropriate response messages based on the recognized emotions and analysis results. This process uses a dialogue generation model such as OpenAI's GPT-3.
[1102] 5. Image generation AI function: Generates images that match the response message and user emotion. For example, DALL·E or similar image generation AI can be used.
[1103] 6. Data transmission function: Sends the generated response message and image URL to the device. Saves the image using a cloud storage service and returns the URL.
[1104] Terminal
[1105] The terminal has the following features:
[1106] 1. Message input function: Providing an interface for users to input messages into a dialogue system, such as a form on a web page or a mobile app.
[1107] 2. Message sending function: Sends the entered message to the server. Data is transferred in JSON format using an HTTP request.
[1108] 3. Data reception function: Receives the response message and image URL sent from the server.
[1109] 4. Display: Displaying the received text and images to the user. This can be achieved using HTML and JavaScript on a web page or mobile app.
[1110] user
[1111] Users interact with the system through an application or web interface. They can type messages and see visual responses and images. For example, if a user types "How are you feeling today?", the server parses the message and generates and sends the appropriate response and associated image.
[1112] Specific examples
[1113] Consider a case where a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. The emotion engine also recognizes the emotion contained in the user's message (for example, anticipation or interest). The response message is generated as "It looks like it's going to be sunny tomorrow. Looking forward to it!" At the same time, an image of a cheerful expression that evokes sunny weather is generated. This response message and the image URL are sent to the user and displayed on the device.
[1114] Prompt Sentence Examples
[1115] Below are some examples of prompts to be used in generative AI models:
[1116] Prompt: "The user asked, 'What's the weather like tomorrow?' They're hoping for sunny weather and would like to see an image of their favorite character with a cheerful expression. Please generate an appropriate response message and image description for this."
[1117] This system allows users to experience emotionally-driven, personalized interactions and visual feedback, enhancing the real-time interaction experience.
[1118] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1119] Step 1:
[1120] The user enters a message into the application or web interface, such as "How are you feeling today?" The input message is stored in text format on the device and sent for further processing.
[1121] Step 2:
[1122] The terminal receives messages entered by the user and sends them to the server through an HTTP request. The messages are converted to JSON format and sent to the server's API endpoint. The input is the user's message, and the output is the JSON data sent to the server.
[1123] Step 3:
[1124] The server receives the message received from the terminal via HTTP request and stores it in the database for analysis. The server is the user's input message and the output is a database storage operation for analysis.
[1125] Step 4:
[1126] The server analyzes the stored messages using natural language processing technology. Specifically, it uses the Google Cloud Natural Language API to understand the meaning and intent of the messages. The input is the user's message read from the database, and the output is the analysis results (e.g., the topic and intent of the message).
[1127] Step 5:
[1128] The server recognizes the user's emotion using an emotion engine (e.g., IBM Watson Tone Analyzer) based on the analysis results. At this stage, the previous analysis results are used as input, and the recognized emotion (e.g., curiosity, anticipation, etc.) is obtained as output.
[1129] Step 6:
[1130] The server uses the analysis results and the recognized emotions as input and generates a response message using a dialogue generation model such as OpenAI GPT-3. The output is the generated response message. For example, a message like "I'm feeling great today. How about you?" is generated.
[1131] Step 7:
[1132] The server generates an image that corresponds to the user's emotion based on the generated response message. It uses image generation AI (e.g., DALL·E) to generate an image that matches the response. The input is the response message and emotion, and the output is the generated image.
[1133] Step 8:
[1134] The server uploads the generated image to an external storage device (e.g., cloud storage) and obtains its URL. The input is the generated image, and the output is the image URL.
[1135] Step 9:
[1136] The server sends the response message and the image URL together in JSON format to the terminal. The input is the response message and the image URL, and the output is the HTTP response sent to the terminal.
[1137] Step 10:
[1138] The terminal analyzes the response message and image URL received from the server and displays them on a web page or application screen. The input is JSON data from the server, and the output is a text message and image displayed to the user.
[1139] (Application example 2)
[1140] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1141] In a dialogue system that allows users to enjoy emotionally rich real-time communication with their favorite characters or celebrities, it is important to accurately recognize the user's emotions and provide response messages and visual feedback that correspond to those emotions. Furthermore, there is a need for a means to efficiently manage the generated images and response messages and provide them to users.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1143] In this invention, the server includes a means for receiving an input message from a user, a means for analyzing the received message and generating a response message based on the content of the message, a means using artificial intelligence for generating an image based on the response message and the user's emotions, and a means for sending the generated response message and image to the user. This allows users to enjoy emotional communication with their favorite characters or celebrities in real time, and enables personalized responses and visual feedback.
[1144] A "user" is an individual who uses the system of the present invention to enjoy conversations with their favorite characters or celebrities.
[1145] An "input message" is a text-based message that a user sends to the system through a terminal.
[1146] The "receiving means" refers to a function and mechanism for capturing an input message sent by a user into the server.
[1147] The "analysis means" refers to a function and mechanism for analyzing a received message using natural language processing technology and understanding its content and the user's intent.
[1148] A "response message" is a text message generated as a reply by the system based on the content analyzed by the analysis means and the user's recognized emotions.
[1149] The "image generation means" refers to a function and mechanism for generating an appropriate image using artificial intelligence based on the response message and the user's emotions.
[1150] "Artificial intelligence" is a technology that uses large amounts of data to analyze and process user emotions and responses, and then generates images and messages accordingly.
[1151] "Transmission means" refers to the functionality and mechanisms for providing the generated response message and image to the user.
[1152] "Real-time" refers to the timeliness in which a response message and image are generated and sent to the user immediately in response to a user's input.
[1153] The present invention relates to a system that allows users to enjoy emotionally rich communication with their favorite characters and celebrities in real time. Specific embodiments of the system are described below.
[1154] System configuration
[1155] server
[1156] The server has the following main functions:
[1157] Message receiving method:
[1158] Receives input messages from users. At this time, the server receives the user's messages in real time via the Internet.
[1159] Analysis method:
[1160] The received message is analyzed using natural language processing technology, such as software like Google BERT and SpaCy, to understand the meaning and intent of the message.
[1161] Emotion recognition means:
[1162] Analyze and recognize emotions from user input messages. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[1163] Response generation method:
[1164] Based on the analysis results and the recognized emotions, an appropriate response message is generated, using a generative AI model such as GPT-3.
[1165] Image generation method:
[1166] Generate images that match the response message and user emotion using image generation AI such as DALL-E or OpenAI CLIP.
[1167] Data transmission method:
[1168] Send the generated text and image URL to the user's device. Upload the image to cloud storage and obtain its URL.
[1169] Terminal
[1170] The terminal performs the following functions:
[1171] Message input method:
[1172] It provides an interface for users to input messages into the dialogue system.
[1173] Message sending method:
[1174] Sends the input message to the server.
[1175] Data receiving method:
[1176] Receive the response message and image URL sent from the server.
[1177] Display means:
[1178] Display the received text and image to the user.
[1179] Operational Overview
[1180] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, where it is analyzed using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[1181] Next, the server uses an emotion recognition means to recognize the emotion from the user's input. In this case, "curiosity" is read from the message. Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[1182] Based on this response message, the image generation means generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message. The device receives the data from the server and displays the text "I'm feeling great today. How about you?" and the image of the favorite character with a cheerful expression to the user.
[1183] Specific examples
[1184] For example, suppose a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. At the same time, the emotion recognition means recognizes the emotion (e.g., anticipation or interest) underlying the user's message. In response, a message such as "It looks like it's going to be sunny tomorrow. Looking forward to it!" is generated. Furthermore, an image of a cheerful expression that evokes the image of sunny weather is generated. This response message and image are sent to the terminal, allowing the user to enjoy it visually.
[1185] This allows the system to allow users to enjoy emotionally rich communication with their favorite characters and celebrities in real time, while providing personalized responses and visual feedback.
[1186] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1187] Step 1:
[1188] The user inputs the message "How are you feeling today?" through the terminal. The terminal receives this input and sends it to the server.
[1189] Input: User input message "How are you feeling today?"
[1190] Output: A request to send a message to the server
[1191] Specific actions: The user enters a message in the device's input interface and presses the send button.
[1192] Step 2:
[1193] The server receives the message sent by the user.
[1194] Input: Message sent from the terminal
[1195] Output: Message data
[1196] Specific operation: The message receiving API on the server receives user messages and stores them in the database.
[1197] Step 3:
[1198] The server parses the received message and performs natural language processing (NLP) to understand its content.
[1199] Input: Received message data
[1200] Output: Analysis results (meaning and intent of the text)
[1201] What it does: Uses an NLP toolkit (e.g. SpaCy, Google BERT) to decompose the message text and analyze its content.
[1202] Step 4:
[1203] The server recognizes the user's emotions from the analysis results and identifies the user's emotions using emotion analysis.
[1204] Input: Analysis results (text meaning and intent)
[1205] Output: Perceived emotion (e.g., curiosity)
[1206] What it does: Analyzes the sentiment of text using a sentiment analysis library (e.g., IBM Watson Tone Analyzer).
[1207] Step 5:
[1208] The server generates a response message based on the analysis results and the recognized emotions.
[1209] Input: Analysis results and recognized emotions
[1210] Output: Response message (e.g. "I'm feeling great today. How about you?")
[1211] Specific operation: Uses response generation AI (e.g., GPT-3) to generate an appropriate response message based on user input.
[1212] Step 6:
[1213] The server uses image generation AI to generate an appropriate image based on the generated response message and the recognized emotion.
[1214] Input: Response message and perceived emotion
[1215] Output: The generated image
[1216] Specific operation: Using image generation AI (e.g., DALL-E, OpenAI CLIP), generate an image corresponding to the response message.
[1217] Step 7:
[1218] The server uploads the generated image to cloud storage and retrieves its URL.
[1219] Input: Generated image
[1220] Output: Image URL
[1221] Specific operation: Calls an API to upload the generated image to cloud storage and obtains the image URL.
[1222] Step 8:
[1223] The server generates a response message and sends the image URL to the terminal.
[1224] Input: Response message and image URL
[1225] Output: Data sent to the terminal
[1226] Specific operation: The server generates a response including a response message and an image URL and sends it to the device.
[1227] Step 9:
[1228] The device receives the data from the server and displays it to the user.
[1229] Input: Response message and image URL sent from the server
[1230] Output: Text and images to display to the user
[1231] Specific operation: The device's receiving interface receives the response from the server and displays text and images on the screen.
[1232] Through these steps, users can enjoy emotionally rich conversations with their favorite characters and celebrities in real time.
[1233] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1234] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1235] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1236] [Fourth embodiment]
[1237] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1238] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1239] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1240] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1241] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1243] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1244] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1245] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1246] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1247] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1248] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1249] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1250] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. This system has the function of receiving an input message from the user, generating a response message based on the content of the message, and generating an image that matches the response message and sending it to the user.
[1251] System configuration
[1252] server
[1253] The server has the following main functions:
[1254] 1. Message receiving function: Receives input messages from the user.
[1255] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[1256] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[1257] 4. Image generation AI function: Generates an image corresponding to the response message.
[1258] 5. Data transmission function: Sends the generated text and image URL to the terminal.
[1259] Terminal
[1260] The terminal performs the following functions:
[1261] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[1262] 2. Message sending function: Sends the input message to the server.
[1263] 3. Data reception function: Receives the response message and image URL sent from the server.
[1264] 4. Display function: displays the received text and images to the user.
[1265] user
[1266] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[1267] Operational Overview
[1268] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its contents using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[1269] Next, the server uses the dialogue model to generate a response message saying, "I'm feeling great today. How about you?" Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message.
[1270] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[1271] Specific examples
[1272] For example, suppose a user asks, "What's the weather like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. In response, it generates a message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It also generates an image of a person with a cheerful expression that evokes sunny weather. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[1273] In this way, this system allows users to enjoy communicating with their favorite idols both verbally and visually.
[1274] The processing flow will be explained below.
[1275] Step 1:
[1276] The user enters a message through the app or web interface.
[1277] Step 2:
[1278] The device receives an input message from the user. For example, the user types, "How are you feeling today?"
[1279] Step 3:
[1280] The device sends the received message to the server, including the message and information such as the user ID.
[1281] Step 4:
[1282] The server receives the incoming message.
[1283] Step 5:
[1284] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[1285] Step 6:
[1286] The server uses the dialogue model to generate a response message based on the analysis results, for example, "I'm feeling great today. How about you?"
[1287] Step 7:
[1288] The server calls the image generation AI based on the generated response message.
[1289] Step 8:
[1290] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[1291] Step 9:
[1292] The server uploads the generated images to an external storage device (such as cloud storage).
[1293] Step 10:
[1294] The server retrieves the image URL from cloud storage.
[1295] Step 11:
[1296] The server sends the generated response message and data including the image URL to the terminal.
[1297] Step 12:
[1298] The terminal receives the response message and image URL received from the server.
[1299] Step 13:
[1300] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[1301] Example 1
[1302] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1303] Conventional dialogue systems simply respond to messages entered by users with text, making it difficult to provide visually appealing, real-time dialogue. This often results in perfunctory conversations without fully understanding the user's emotions and intentions. Furthermore, conventional systems are limited in the images they generate as responses, making it impossible to generate images that reflect the diverse facial expressions and postures expected by users.
[1304] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1305] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for sending the generated response message and images to the user, means for uploading the images to data storage and sending their URLs to the user, and means for analyzing the message using a natural language processing model and generating images using a generative AI model. This allows the user to receive a response including images with a variety of facial expressions and poses that reflect their emotions and intentions, enabling them to enjoy visually rich, real-time dialogue.
[1306] An "input message" is text data that a user sends through interaction with the system.
[1307] "Analysis" refers to the process performed to understand the meaning and intent of a received message.
[1308] A "response message" is text data that is generated as a response to an input message by the system.
[1309] "Image generation" refers to the process of generating images with specific facial expressions and poses based on text data and analysis results.
[1310] "Data storage" refers to external storage devices and cloud services for saving generated images and data.
[1311] "URL" means a uniform resource identifier for locating resources on the Internet.
[1312] A "natural language processing model" is an algorithm or program that allows a computer to understand and generate human language.
[1313] A "generative AI model" is an artificial intelligence model that generates new data based on input prompts.
[1314] A "server" is a computer system that processes user requests and provides various functions of an interactive system.
[1315] A "terminal" is a device used by a user (e.g., a smartphone or PC) that provides an interface for interacting with the system.
[1316] This invention relates to a system that allows users to enjoy real-time conversations with their favorite characters and celebrities. The system generates appropriate response messages and related images in response to messages entered by users and provides them to the users. This system is primarily composed of a server and a terminal.
[1317] Server configuration and functions
[1318] Message receiving function
[1319] The server has the ability to receive input messages from the user, which can be done using a standard HTTP POST request.
[1320] Message analysis function using natural language processing
[1321] The server analyzes the received message using a natural language processing model (e.g., Hugging Face Transformers) to understand the content and intent of the message.
[1322] Response message generation function
[1323] Based on the analysis results, the server uses a dialogue model (e.g., OpenAI's GPT-3) to generate an appropriate response message, providing an appropriate response to the user's question or comment.
[1324] Image generation function
[1325] Based on the generated response message, an image generation AI (e.g., DALL-E) is used to generate images with different facial expressions and poses, which are then uploaded to cloud storage (e.g., Amazon S3).
[1326] Data transmission function
[1327] The generated response message and image URL are sent to the user's device as a JSON response, allowing the user to receive both the text and the image.
[1328] Device configuration and functions
[1329] Message input function
[1330] The device used by the user (e.g., a smartphone or a computer) provides an interface in which the message can be entered, which includes a text box.
[1331] Message sending function
[1332] The device has the ability to send messages entered by the user to a server, typically using an HTTP POST request.
[1333] Data reception function
[1334] It has the function of receiving response messages and image URLs from the server, and the received data is processed through the application or web browser.
[1335] Display function
[1336] The device displays received text messages and images to the user, allowing the user to enjoy intuitive interaction.
[1337] Specific examples
[1338] For example, here is what happens when a user types "What's the weather like tomorrow?" When a user types "What's the weather like tomorrow?" into a smartphone app, the message is sent to the server via an HTTP POST request. The server analyzes the message using Hugging Face Transformers and recognizes it as a "weather-related question." Next, the server uses OpenAI's GPT-3 to generate a response message saying, "It looks like it's going to be sunny tomorrow. Looking forward to it!" It then uses image generation AI (DALL-E) to generate an image of a "cheerful expression that evokes sunny weather" and uploads it to Amazon S3. Finally, the server sends the response message and the image URL to the user's device, where the user can view it.
[1339] Prompt Sentence Examples
[1340] Here are some example prompts for a generative AI model:
[1341] Natural Language Parsing Prompt:
[1342] User message: "What's the weather like tomorrow?"
[1343] Analysis results: Weather questions
[1344] Dialogue model prompt:
[1345] User message: "What's the weather like tomorrow?"
[1346] Model's response: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[1347] Image generation AI prompt:
[1348] Response message: "It looks like it's going to be sunny tomorrow. Looking forward to it!"
[1349] Generated image: A character image with a cheerful expression that evokes a sunny day
[1350] By combining these functions and processes, the present invention realizes a system that allows users to simultaneously enjoy conversations that reflect their emotions and intentions and visual feedback.
[1351] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1352] Step 1: User message input
[1353] The user inputs a conversation message into the system through a terminal application or a web interface. The input data is a text message, such as "How are you feeling today?" The input message is stored in memory.
[1354] Step 2: Send the message to the server
[1355] The terminal sends the message entered by the user to the server. Specifically, it uses an HTTP POST request to send the entered text message to the server in JSON format. The input data is the user's message, and the output is an HTTP request in JSON format.
[1356] Step 3: Receiving the message on the server
[1357] The server receives messages sent from the device, for example, a request at a specific endpoint using a web framework like Flask. The input data is the request in JSON format, and the output is a text message extracted from that request.
[1358] Step 4: Message analysis using natural language processing
[1359] The server inputs the received message into a natural language processing model (for example, Hugging Face Transformers). The model analyzes the content of the text message and generates data to understand its meaning and intent. The input data is the user's text message, and the output data is the information on intent and meaning as a result of the analysis. Specifically, it analyzes the message "How are you feeling today?" and classifies it as a "question about feelings."
[1360] Step 5: Generate a response message
[1361] The server uses a dialogue model (e.g., OpenAI's GPT-3) based on the analysis results to generate an appropriate response message. The input data is the message analysis result, and the output data is the response message. For example, a response message such as "I'm feeling great today. How about you?" is generated.
[1362] Step 6: Image generation
[1363] Based on the generated response message, the server uses an image generation AI (e.g., DALL-E) to generate a corresponding image. The input data is the response message, and the output data is the generated image. Specifically, an image of a "cheerful expression" is generated and uploaded to cloud storage.
[1364] Step 7: Uploading images to data storage
[1365] The server uploads the generated image to data storage (e.g., Amazon S3) and retrieves its URL. The input data is the generated image, and the output data is the image URL. This allows for long-term storage and access of the image.
[1366] Step 8: Sending response data from the server to the device
[1367] The server sends the generated response message and the image URL together in JSON format to the terminal. The input data is the response message and the image URL, and the output data is the HTTP response in JSON format.
[1368] Step 9: Receive and display data on your device
[1369] The device receives the data sent from the server and displays it to the user. Specifically, it displays the text message "I'm feeling great today. How about you?" and an image of a cheerful expression on the app screen. The input data is the JSON response from the server, and the output data is the text and image displayed to the user.
[1370] In this way, users can enjoy interacting visually and intuitively.
[1371] (Application example 1)
[1372] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1373] Traditional online shopping lacks a way to seamlessly interact with favorite characters or celebrities, a new and unprecedented user experience. This results in a monotonous product browsing experience for users, which reduces the appeal of shopping itself. Furthermore, the lack of real-time customer support and product introductions reduces user choice and satisfaction.
[1374] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1375] In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message and generating a response message based on the content of the message, means for generating images with different facial expressions and poses based on the response message, means for transmitting the generated response message and image to the user, means for responding in real time to the user's questions and requests so that the user can be guided through the virtual environment, and means for displaying images and information related to the responses. This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and allows them to obtain detailed information about products in real time.
[1376] "Users" refer to people who use the system to shop while enjoying conversations with their favorite characters and celebrities.
[1377] "Input Message" refers to a message such as a question or request that a User sends through the System.
[1378] "Means for receiving" refers to the function of the system to receive an input message sent by a user and send it to a server.
[1379] "Means of analysis" refers to algorithms or functions that understand the received input message and analyze its meaning and intent.
[1380] "Response message" refers to a reply message generated by the system based on the analysis results.
[1381] "Means of generation" refers to the function of using AI or other means to create images with different facial expressions and postures related to the response message.
[1382] "Means for transmitting" refers to a function for transmitting the generated response message and image to provide them to the user.
[1383] "Virtual Environment" refers to the virtual space in which users shop online.
[1384] "Means of responding in real time" refers to the ability to generate a response to user input immediately and engage in dialogue.
[1385] "Means for displaying images and information" refers to the function of visually displaying the generated images and related information to the user.
[1386] MODE FOR CARRYING OUT THE INVENTION
[1387] A specific embodiment of the present invention will now be described. The system of the present invention allows users to enjoy shopping while interacting with their favorite characters or celebrities in a virtual environment. The main configuration and functions of the system are shown below.
[1388] System configuration
[1389] server
[1390] The server has the following main functions:
[1391] 1. Message receiving function: Receives input messages from the user.
[1392] 2. Content analysis function: Analyzes received messages and understands their meaning and intent.
[1393] 3. Response generation function: Generates an appropriate response message based on the analysis results.
[1394] 4. Image generation AI function: Generates images with different facial expressions and postures corresponding to the response message.
[1395] 5. Data transmission function: Sends the generated text and image URL to the user's device.
[1396] Terminal
[1397] The terminal performs the following functions:
[1398] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[1399] 2. Message sending function: Sends the entered message to the server.
[1400] 3. Data receiving function: Receives the response message and image URL sent from the server.
[1401] 4. Display function: Display the received text and images to the user.
[1402] Processing flow
[1403] The user types a message through the terminal, asking a specific question such as "What are the features of this product?" This generates a prompt like the following:
[1404] Example prompt: What are the features of this product?
[1405] The server receives the message and analyzes it using the content analysis function. Based on the analysis results, the response generation function generates a response message such as "This product has plenty of storage space and a simple, stylish design."
[1406] Next, the image generation AI function generates an image related to the response message (for example, a multi-angle photo of a product) and uploads it to cloud storage. A response message containing the URL of the generated image is sent from the server to the user's device.
[1407] The device receives data from the server and displays an image of the product along with text such as, "This product has ample storage space and a simple, stylish design." This allows users to enjoy shopping while interacting with their favorite characters or celebrities, and obtain detailed information about the product in real time.
[1408] Hardware and software used
[1409] The following hardware and software are used to implement this system:
[1410] Hardware: Smartphone (iOS, Android), Server (Cloud-based)
[1411] Software: Python, requests library, Flask (for server-side implementation)
[1412] This allows users to shop in a virtual space while enjoying real-time interactions with their favorite characters and celebrities.
[1413] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1414] Step 1:
[1415] The user inputs a message through the terminal. The input message is a question or request that the user makes to their favorite character or celebrity. For example, a specific question such as "What are the features of this product?" The input message is sent to the server using the message input interface.
[1416] Step 2:
[1417] The server analyzes the received message. It uses content analysis to understand the meaning and intent of the message. A natural language processing (NLP) algorithm is used for this analysis. The message from the user, "What are the features of this product?", is passed as input, and the analysis results are obtained as output. The analysis results include information such as the message being a "question about the product's features."
[1418] Step 3:
[1419] The server generates a response message based on the analysis results. An appropriate reply is generated using the response generation function. For example, based on the analysis results, a response message such as "This product has plenty of storage space and a simple, stylish design" is generated. The input is the analysis results, and the output is the generated response message.
[1420] Step 4:
[1421] The server generates an image based on the generated response message. This process utilizes image generation AI functions. The response message is input as a prompt into the AI model, which then generates an appropriate image (for example, a multi-angle photo of a product). The response message is provided as input, and the generated image is obtained as output.
[1422] Step 5:
[1423] The server uploads the generated image to cloud storage. A URL for the uploaded image is generated and sent to the user device along with a response message. The input is the generated image, and the output is data containing the image URL and the response message.
[1424] Step 6:
[1425] The terminal receives the data sent from the server. Using the data reception function, the response message and image URL are received. The input is the data from the server (the response message and URL), and the output is the received data.
[1426] Step 7:
[1427] The device displays the received response message and image to the user. Using the display function, the text and image are displayed on the user's screen. The received data is the input, and a visual display is the output. This allows the user to enjoy the experience of shopping while interacting with their favorite character or celebrity.
[1428] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1429] The present invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities, and in particular, is capable of recognizing the user's emotions and reflecting them in the content of the interaction and image generation. This system has the function of receiving an input message from the user, generating a response message based on the content and the user's emotions, and further generating an image corresponding to the response message and sending it to the user.
[1430] System configuration
[1431] server
[1432] The server has the following main functions:
[1433] 1. Message receiving function: Receives input messages from the user.
[1434] 2. Content analysis function: Analyzes received messages using natural language processing technology to understand their meaning and intent.
[1435] 3. Emotion recognition function (emotion engine): Analyzes and recognizes the user's emotions from the user's input message.
[1436] 4. Response generation function: Generates an appropriate response message based on the analysis results and the recognized emotions.
[1437] 5. Image generation AI function: Generates images that match the response message and user emotions.
[1438] 6. Data transmission function: Sends the generated text and image URL to the terminal.
[1439] Terminal
[1440] The terminal performs the following functions:
[1441] 1. Message input function: Provides an interface for users to input messages into the dialogue system.
[1442] 2. Message sending function: Sends the input message to the server.
[1443] 3. Data reception function: Receives the response message and image URL sent from the server.
[1444] 4. Display function: displays the received text and images to the user.
[1445] user
[1446] Users interact with the system through an app or web interface by typing messages and visually seeing responses.
[1447] Operational Overview
[1448] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, which receives the message and analyzes its content using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[1449] Next, the server uses an emotion engine to recognize the emotion from the user's input. In this case, the message reads "curiosity." Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[1450] Based on this response message, the image generation AI generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and the URL is sent to the device along with the response message.
[1451] The device receives the data from the server and displays to the user an image of their favorite idol with a cheerful expression along with the text "I'm feeling great today. How about you?", giving the user the experience of seeing their favorite idol actually converse and express their emotions.
[1452] Specific examples
[1453] For example, if a user asks, "What's the weather going to be like tomorrow?", the server analyzes the message and recognizes that it is a weather-related question. At the same time, the emotion engine recognizes the emotion (e.g., expectation or interest) underlying the user's message.
[1454] In response, a message is generated saying, "It looks like it's going to be sunny tomorrow. I'm looking forward to it!" In addition, an image of a cheerful expression that evokes sunny weather is generated. This response message and image are sent to the device, allowing the user to enjoy the visuals.
[1455] In this way, the system allows users to enjoy both verbal and visual communication with their favorite idols. In particular, the addition of emotion recognition functionality provides more personalized responses, improving the user experience.
[1456] The processing flow will be explained below.
[1457] Step 1:
[1458] The user types a message through the app or web interface, for example, "How are you feeling today?"
[1459] Step 2:
[1460] The terminal receives an input message from the user.
[1461] Step 3:
[1462] The device sends the received message to the server, including the message and information such as the user ID.
[1463] Step 4:
[1464] The server receives the incoming message.
[1465] Step 5:
[1466] The server analyzes the received message using natural language processing (NLP). This analysis allows it to understand the meaning and intent of the message. For example, it can determine that a message like "How are you feeling today?" is a "question about feelings."
[1467] Step 6:
[1468] The server uses an emotion engine to recognize the user's emotion from the user's input message, for example, recognizing the user's curiosity from the message "How are you feeling today?"
[1469] Step 7:
[1470] The server uses the dialogue model to generate an appropriate response message based on the analysis results and the recognized emotions, for example, "I'm feeling great today. How about you?"
[1471] Step 8:
[1472] The server calls the image generation AI based on the generated response message.
[1473] Step 9:
[1474] The server uses image generation AI to generate an image that corresponds to the response message. For example, an image with a "cheerful expression" is generated.
[1475] Step 10:
[1476] The server uploads the generated images to an external storage device (cloud storage).
[1477] Step 11:
[1478] The server retrieves the image URL from cloud storage.
[1479] Step 12:
[1480] The server sends the generated response message and data including the image URL to the terminal.
[1481] Step 13:
[1482] The terminal receives the response message and image URL received from the server.
[1483] Step 14:
[1484] The device will display the received response message and image to the user. For example, the text "I'm feeling great today. How about you?" will be displayed along with an image of your favorite character with a cheerful expression.
[1485] Example 2
[1486] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1487] Today's users desire real-time interactive experiences with their favorite characters and celebrities, but conventional dialogue systems lack emotional response and visual feedback. Conventional systems struggle to automatically generate personalized responses and visual expressions that reflect the user's emotions, and improvements to the user experience are needed.
[1488] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for analyzing the received message using natural language processing technology and understanding its content, means for recognizing the user's emotion based on the analysis result, means for generating a response message based on the recognized emotion and the analysis result, means for generating an image that matches the emotion based on the response message, and means for sending the URL of the generated response message and image to the user. This allows the user to receive personalized dialogue and visual feedback according to their emotion, improving their real-time dialogue experience.
[1489] A "user" is an individual who uses a system and seeks an interactive experience.
[1490] An "input message" is information in the form of a string that a user sends to the system.
[1491] A "server" is a computer system that receives messages from users, analyzes them, recognizes emotions, generates responses, and generates images.
[1492] "Natural language processing technology" refers to a set of techniques and algorithms that enable computers to understand and analyze human language.
[1493] The "emotion engine" is a software module for recognizing and classifying emotions from user input messages.
[1494] The "response message" is information in the form of a string that the server generates based on the analysis results and the recognized emotion.
[1495] "Image generation AI" is an artificial intelligence technology that generates visual images based on input text and conditions.
[1496] "External storage device" refers to a remote storage service for saving generated data (e.g., image files).
[1497] A "URL" is an address notation for specifying resources on the Internet, and is a type of URI (Uniform Resource Identifier) that indicates the location where the generated image is saved.
[1498] "Means for sending to the user" refers to the communication means for transferring the generated response message and image URL to the user's terminal.
[1499] This invention relates to a system that allows users to enjoy real-time interactions with their favorite characters and celebrities. Specifically, it has the function of receiving messages input by the user, generating response messages and images based on the content and emotions of the messages, and sending them to the user. This system is mainly composed of three main components: a server, a terminal, and the user.
[1500] System configuration
[1501] server
[1502] The server has the following main functions:
[1503] 1. Message receiving function: Receives input messages from the user. This function can be implemented using, for example, the Python Flask framework.
[1504] 2. Content analysis function: Analyzes received messages using natural language processing technology (e.g., Google Cloud Natural Language API) to understand their meaning.
[1505] 3. Emotion recognition function (emotion engine): Recognizes emotions based on the analyzed content. This can be implemented, for example, using IBM Watson Tone Analyzer.
[1506] 4. Response generation function: Generates appropriate response messages based on the recognized emotions and analysis results. This process uses a dialogue generation model such as OpenAI's GPT-3.
[1507] 5. Image generation AI function: Generates images that match the response message and user emotion. For example, DALL·E or similar image generation AI can be used.
[1508] 6. Data transmission function: Sends the generated response message and image URL to the device. Saves the image using a cloud storage service and returns the URL.
[1509] Terminal
[1510] The terminal has the following features:
[1511] 1. Message input function: Providing an interface for users to input messages into a dialogue system, such as a form on a web page or a mobile app.
[1512] 2. Message sending function: Sends the entered message to the server. Data is transferred in JSON format using an HTTP request.
[1513] 3. Data reception function: Receives the response message and image URL sent from the server.
[1514] 4. Display: Displaying the received text and images to the user. This can be achieved using HTML and JavaScript on a web page or mobile app.
[1515] user
[1516] Users interact with the system through an application or web interface. They can type messages and see visual responses and images. For example, if a user types "How are you feeling today?", the server parses the message and generates and sends the appropriate response and associated image.
[1517] Specific examples
[1518] Consider a case where a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. The emotion engine also recognizes the emotion contained in the user's message (for example, anticipation or interest). The response message is generated as "It looks like it's going to be sunny tomorrow. Looking forward to it!" At the same time, an image of a cheerful expression that evokes sunny weather is generated. This response message and the image URL are sent to the user and displayed on the device.
[1519] Prompt Sentence Examples
[1520] Below are some examples of prompts to be used in generative AI models:
[1521] Prompt: "The user asked, 'What's the weather like tomorrow?' They're hoping for sunny weather and would like to see an image of their favorite character with a cheerful expression. Please generate an appropriate response message and image description for this."
[1522] This system allows users to experience emotionally-driven, personalized interactions and visual feedback, enhancing the real-time interaction experience.
[1523] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1524] Step 1:
[1525] The user enters a message into the application or web interface, such as "How are you feeling today?" The input message is stored in text format on the device and sent for further processing.
[1526] Step 2:
[1527] The terminal receives messages entered by the user and sends them to the server through an HTTP request. The messages are converted to JSON format and sent to the server's API endpoint. The input is the user's message, and the output is the JSON data sent to the server.
[1528] Step 3:
[1529] The server receives the message received from the terminal via HTTP request and stores it in the database for analysis. The server is the user's input message and the output is a database storage operation for analysis.
[1530] Step 4:
[1531] The server analyzes the stored messages using natural language processing technology. Specifically, it uses the Google Cloud Natural Language API to understand the meaning and intent of the messages. The input is the user's message read from the database, and the output is the analysis results (e.g., the topic and intent of the message).
[1532] Step 5:
[1533] The server recognizes the user's emotion using an emotion engine (e.g., IBM Watson Tone Analyzer) based on the analysis results. At this stage, the previous analysis results are used as input, and the recognized emotion (e.g., curiosity, anticipation, etc.) is obtained as output.
[1534] Step 6:
[1535] The server uses the analysis results and the recognized emotions as input and generates a response message using a dialogue generation model such as OpenAI GPT-3. The output is the generated response message. For example, a message like "I'm feeling great today. How about you?" is generated.
[1536] Step 7:
[1537] The server generates an image that corresponds to the user's emotion based on the generated response message. It uses image generation AI (e.g., DALL·E) to generate an image that matches the response. The input is the response message and emotion, and the output is the generated image.
[1538] Step 8:
[1539] The server uploads the generated image to an external storage device (e.g., cloud storage) and obtains its URL. The input is the generated image, and the output is the image URL.
[1540] Step 9:
[1541] The server sends the response message and the image URL together in JSON format to the terminal. The input is the response message and the image URL, and the output is the HTTP response sent to the terminal.
[1542] Step 10:
[1543] The terminal analyzes the response message and image URL received from the server and displays them on a web page or application screen. The input is JSON data from the server, and the output is a text message and image displayed to the user.
[1544] (Application example 2)
[1545] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1546] In a dialogue system that allows users to enjoy emotionally rich real-time communication with their favorite characters or celebrities, it is important to accurately recognize the user's emotions and provide response messages and visual feedback that correspond to those emotions. Furthermore, there is a need for a means to efficiently manage the generated images and response messages and provide them to users.
[1547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1548] In this invention, the server includes a means for receiving an input message from a user, a means for analyzing the received message and generating a response message based on the content of the message, a means using artificial intelligence for generating an image based on the response message and the user's emotions, and a means for sending the generated response message and image to the user. This allows users to enjoy emotional communication with their favorite characters or celebrities in real time, and enables personalized responses and visual feedback.
[1549] A "user" is an individual who uses the system of the present invention to enjoy conversations with their favorite characters or celebrities.
[1550] An "input message" is a text-based message that a user sends to the system through a terminal.
[1551] The "receiving means" refers to a function and mechanism for capturing an input message sent by a user into the server.
[1552] The "analysis means" refers to a function and mechanism for analyzing a received message using natural language processing technology and understanding its content and the user's intent.
[1553] A "response message" is a text message generated as a reply by the system based on the content analyzed by the analysis means and the user's recognized emotions.
[1554] The "image generation means" refers to a function and mechanism for generating an appropriate image using artificial intelligence based on the response message and the user's emotions.
[1555] "Artificial intelligence" is a technology that uses large amounts of data to analyze and process user emotions and responses, and then generates images and messages accordingly.
[1556] "Transmission means" refers to the functionality and mechanisms for providing the generated response message and image to the user.
[1557] "Real-time" refers to the timeliness in which a response message and image are generated and sent to the user immediately in response to a user's input.
[1558] The present invention relates to a system that allows users to enjoy emotionally rich communication with their favorite characters and celebrities in real time. Specific embodiments of the system are described below.
[1559] System configuration
[1560] server
[1561] The server has the following main functions:
[1562] Message receiving method:
[1563] Receives input messages from users. At this time, the server receives the user's messages in real time via the Internet.
[1564] Analysis method:
[1565] The received message is analyzed using natural language processing technology, such as software like Google BERT and SpaCy, to understand the meaning and intent of the message.
[1566] Emotion recognition means:
[1567] Analyze and recognize emotions from user input messages. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[1568] Response generation method:
[1569] Based on the analysis results and the recognized emotions, an appropriate response message is generated, using a generative AI model such as GPT-3.
[1570] Image generation method:
[1571] Generate images that match the response message and user emotion using image generation AI such as DALL-E or OpenAI CLIP.
[1572] Data transmission method:
[1573] Send the generated text and image URL to the user's device. Upload the image to cloud storage and obtain its URL.
[1574] Terminal
[1575] The terminal performs the following functions:
[1576] Message input method:
[1577] It provides an interface for users to input messages into the dialogue system.
[1578] Message sending method:
[1579] Sends the input message to the server.
[1580] Data receiving method:
[1581] Receive the response message and image URL sent from the server.
[1582] Display means:
[1583] Display the received text and image to the user.
[1584] Operational Overview
[1585] The user types a message such as "How are you feeling today?" into the device. This message is instantly sent to the server, where it is analyzed using a natural language processing algorithm. As a result of the analysis, the message is classified as a "question about feelings."
[1586] Next, the server uses an emotion recognition means to recognize the emotion from the user's input. In this case, "curiosity" is read from the message. Based on the analysis results and the recognized emotion, the server uses a dialogue model to generate a response message such as "I'm feeling great today. How about you?"
[1587] Based on this response message, the image generation means generates an image of the favorite character with a "cheerful expression." This image is uploaded to cloud storage, and its URL is sent to the device along with the response message. The device receives the data from the server and displays the text "I'm feeling great today. How about you?" and the image of the favorite character with a cheerful expression to the user.
[1588] Specific examples
[1589] For example, suppose a user asks, "What's the weather going to be like tomorrow?" In this case, the server analyzes the message and recognizes that it is a question about the weather. At the same time, the emotion recognition means recognizes the emotion (e.g., anticipation or interest) underlying the user's message. In response, a message such as "It looks like it's going to be sunny tomorrow. Looking forward to it!" is generated. Furthermore, an image of a cheerful expression that evokes the image of sunny weather is generated. This response message and image are sent to the terminal, allowing the user to enjoy it visually.
[1590] This allows the system to allow users to enjoy emotionally rich communication with their favorite characters and celebrities in real time, while providing personalized responses and visual feedback.
[1591] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1592] Step 1:
[1593] The user inputs the message "How are you feeling today?" through the terminal. The terminal receives this input and sends it to the server.
[1594] Input: User input message "How are you feeling today?"
[1595] Output: A request to send a message to the server
[1596] Specific actions: The user enters a message in the device's input interface and presses the send button.
[1597] Step 2:
[1598] The server receives the message sent by the user.
[1599] Input: Message sent from the terminal
[1600] Output: Message data
[1601] Specific operation: The message receiving API on the server receives user messages and stores them in the database.
[1602] Step 3:
[1603] The server parses the received message and performs natural language processing (NLP) to understand its content.
[1604] Input: Received message data
[1605] Output: Analysis results (meaning and intent of the text)
[1606] What it does: Uses an NLP toolkit (e.g. SpaCy, Google BERT) to decompose the message text and analyze its content.
[1607] Step 4:
[1608] The server recognizes the user's emotions from the analysis results and identifies the user's emotions using emotion analysis.
[1609] Input: Analysis results (text meaning and intent)
[1610] Output: Perceived emotion (e.g., curiosity)
[1611] What it does: Analyzes the sentiment of text using a sentiment analysis library (e.g., IBM Watson Tone Analyzer).
[1612] Step 5:
[1613] The server generates a response message based on the analysis results and the recognized emotions.
[1614] Input: Analysis results and recognized emotions
[1615] Output: Response message (e.g. "I'm feeling great today. How about you?")
[1616] Specific operation: Uses response generation AI (e.g., GPT-3) to generate an appropriate response message based on user input.
[1617] Step 6:
[1618] The server uses image generation AI to generate an appropriate image based on the generated response message and the recognized emotion.
[1619] Input: Response message and perceived emotion
[1620] Output: The generated image
[1621] Specific operation: Using image generation AI (e.g., DALL-E, OpenAI CLIP), generate an image corresponding to the response message.
[1622] Step 7:
[1623] The server uploads the generated image to cloud storage and retrieves its URL.
[1624] Input: Generated image
[1625] Output: Image URL
[1626] Specific operation: Calls an API to upload the generated image to cloud storage and obtains the image URL.
[1627] Step 8:
[1628] The server generates a response message and sends the image URL to the terminal.
[1629] Input: Response message and image URL
[1630] Output: Data sent to the terminal
[1631] Specific operation: The server generates a response including a response message and an image URL and sends it to the device.
[1632] Step 9:
[1633] The device receives the data from the server and displays it to the user.
[1634] Input: Response message and image URL sent from the server
[1635] Output: Text and images to display to the user
[1636] Specific operation: The device's receiving interface receives the response from the server and displays text and images on the screen.
[1637] Through these steps, users can enjoy emotionally rich conversations with their favorite characters and celebrities in real time.
[1638] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1639] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1640] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1641] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1642] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1643] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1644] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1645] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1646] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1647] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1648] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1649] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1650] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1651] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1652] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1653] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1654] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1655] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1656] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1657] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1658] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1659] The following is further disclosed regarding the above embodiment.
[1660] (Claim 1)
[1661] means for receiving input messages from a user;
[1662] means for analyzing the received message and generating a response message based on the content of the message;
[1663] means for generating images with different facial expressions and postures based on the response message;
[1664] means for transmitting the generated response message and image to a user;
[1665] A system including:
[1666] (Claim 2)
[1667] 10. The system of claim 1, further comprising means for generating a response message based on the intent or sentiment of the analyzed message.
[1668] (Claim 3)
[1669] 10. The system of claim 1, further comprising means for uploading the generated image to an external storage device and transmitting the URL thereof to the user.
[1670] "Example 1"
[1671] (Claim 1)
[1672] means for receiving input messages from a user;
[1673] means for analyzing the received message and generating a response message based on the content of the message;
[1674] means for generating images with different facial expressions and postures based on the response message;
[1675] means for transmitting the generated response message and image to a user;
[1676] means for uploading said image to a data storage and sending the URL thereof to the user;
[1677] A system including:
[1678] (Claim 2)
[1679] 10. The system of claim 1, further comprising means for generating a response message based on the intent or sentiment of the analyzed message.
[1680] (Claim 3)
[1681] 10. The system of claim 1, further comprising means for analyzing the message using a natural language processing model and generating the image using a generative AI model.
[1682] "Application Example 1"
[1683] (Claim 1)
[1684] means for receiving input messages from a user;
[1685] means for analyzing the received message and generating a response message based on the content of the message;
[1686] means for generating images with different facial expressions and postures based on the response message;
[1687] means for transmitting the generated response message and image to a user;
[1688] a means for responding in real time to the user's questions and requests in order to guide the user through the virtual environment; and
[1689] means for displaying images or information related to said responses;
[1690] A system including:
[1691] (Claim 2)
[1692] 10. The system of claim 1, further comprising means for generating a response message based on the intent or sentiment of the analyzed message.
[1693] (Claim 3)
[1694] 10. The system of claim 1, further comprising means for uploading the generated image to an external storage device and transmitting the URL thereof to the user.
[1695] "Example 2: Combining Emotion Engines"
[1696] (Claim 1)
[1697] means for receiving input messages from a user;
[1698] means for analyzing the received message using natural language processing technology and understanding its contents;
[1699] A means for recognizing user emotions based on the analysis results;
[1700] means for generating a response message based on the recognized emotion and the analysis result;
[1701] means for generating an image that matches the emotion based on the response message;
[1702] A means for sending the generated response message and image URL to the user;
[1703] A system including:
[1704] (Claim 2)
[1705] 10. The system of claim 1, further comprising means for generating a response message based on the intent and perceived sentiment of the analyzed message.
[1706] (Claim 3)
[1707] 10. The system of claim 1, further comprising means for uploading the generated image to an external storage device and transmitting the URL thereof to the user.
[1708] "Application example 2 when combining emotion engines"
[1709] (Claim 1)
[1710] means for receiving input messages from a user;
[1711] means for analyzing the received message and generating a response message based on the content of the message;
[1712] an artificial intelligence-based means for generating an image based on the response message and the user's emotion;
[1713] means for transmitting the generated response message and image to a user;
[1714] A system including:
[1715] (Claim 2)
[1716] 10. The system of claim 1, further comprising means for generating a response message and image based on the intent and sentiment of the analyzed message.
[1717] (Claim 3)
[1718] 10. The system of claim 1, further comprising means for uploading the generated image to an external storage device and transmitting the URL thereof to the user. [Explanation of symbols]
[1719] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving input messages from a user; means for analyzing the received message and generating a response message based on the content of the message; means for generating images with different facial expressions and postures based on the response message; means for transmitting the generated response message and image to a user; A system including:
2. 10. The system of claim 1, further comprising means for generating a response message based on the intent or sentiment of the analyzed message.
3. 2. The system according to claim 1, further comprising means for uploading the generated image to an external storage device and transmitting the URL thereof to the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A