System
The system addresses the lack of personalized empathetic dialogue by generating avatars and engaging in emotionally intelligent conversations, reducing loneliness and improving psychological stability.
Patent Information
- Application Number
- JP2024138593
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Current systems lack the capability to provide sustained, personalized, and empathetic dialogue for individuals feeling lonely, particularly the elderly and those with busy lives, leading to psychological and physical health risks.
A system that includes inputting basic user information, generating an avatar using AI, engaging in dialogue through a conversational AI, analyzing emotions, generating empathetic responses, storing dialogue and emotional information, and dynamically updating dialogue content based on user changes in interests and hobbies, while ensuring secure data storage.
Provides personalized and empathetic dialogue, reducing feelings of loneliness and achieving psychological stability by adapting to users' evolving interests and emotions.
Smart Images

Figure 2026036078000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the number of people feeling lonely and isolated is increasing. This is particularly concerning for the elderly and those living busy lives, as they have less interaction with family and friends and a lack of emotional support. This sense of loneliness can lead to psychological problems and even physical health risks. Therefore, systems capable of sustained, empathetic dialogue with these lonely people are needed. However, current technology lacks a sufficient number of systems capable of providing appropriate, personalized dialogue for each user. This makes it difficult to alleviate loneliness and achieve psychological stability for users. [Means for solving the problem]
[0005] The present invention provides a system including a means for inputting basic information from a user, a means for generating an avatar using an image generation artificial intelligence (AI) based on the input basic information, and a means for displaying the generated avatar to the user. The system also includes a means for activating a conversation generation AI based on the user's input and engaging in a dialogue, and a means for analyzing emotions from the user's input during the dialogue. The system further includes a means for generating an empathetic response based on the analyzed emotional information, a means for storing the generated dialogue and emotional information in a database, and a means for updating the content of the next dialogue based on the stored information. The system also includes a means for encrypting the user's personal information and securely storing the avatar and its data, a means for learning changes in the user's hobbies and interests, and optimizing the dialogue content based on the learning, and a means for dynamically updating the dialogue content in response to new hobbies and interests. This system provides users with personalized, empathetic, and sustained dialogue, reducing feelings of loneliness and achieving psychological stability.
[0006] A "user" is an individual who utilizes the system to enter basic information and interact with an avatar.
[0007] "Basic information" refers to data such as the user's name, age, hobbies and interests that the user provides to the system.
[0008] "Image generation artificial intelligence" is artificial intelligence that generates an avatar corresponding to a user based on input basic information.
[0009] An "avatar" is a personified character generated by image generation artificial intelligence that empathetically interacts with the user.
[0010] "Conversational generative artificial intelligence" is artificial intelligence that generates appropriate dialogue content based on user input and communicates with the user.
[0011] "Sentiment analysis" is the process of extracting emotional signals from user input and recognizing that emotion.
[0012] "Empathetic response" is a process of generating a response that is appropriate to the user's emotions based on the results of emotion analysis.
[0013] The "database" is a digital storage system for storing the generated conversation and emotion information for later reference.
[0014] "Encryption" is a technology that converts data using a specific algorithm and stores it securely in order to protect users' personal information.
[0015] "Dynamic updating" is the process of changing the dialogue in real time in response to a user's new hobbies and interests.
[0016] "Optimization" is the process of improving the interaction and providing a more personalized experience in response to changes in the user's tastes and interests. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention is a system that provides personalized, empathetic dialogue to users who feel lonely. This system uses artificial intelligence to prompt users to enter basic information and engage in dialogue through a generated avatar. It also analyzes the user's emotions and generates responses based on those emotions. The program for this system is described in detail below.
[0039] Initial Setup and Avatar Creation
[0040] First, a user accesses the platform and enters basic information (name, age, hobbies and interests). This basic information is sent from the device to the server, which then creates a user profile based on it. Next, the server uses image generation artificial intelligence to generate an avatar for the user. This avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[0041] Conversation initiation and emotional understanding
[0042] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" can be generated. In response, the user inputs text or voice, and the device sends the input to the server.
[0043] The server analyzes this input using an emotion analysis module to extract the user's emotions (e.g., interest, joy). It then generates an empathetic response based on the analyzed emotions. For example, if a user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via their device.
[0044] Continuous personalization
[0045] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This allows the server to have basic data to better personalize the content of the next conversation. For example, if the user talks about a new hobby (e.g., painting), this information will be reflected in future conversations. The conversation can continue in the form of, "Regarding the painting we talked about last time, have you painted any works recently?"
[0046] Safety and Privacy Controls
[0047] Users' personal information is encrypted and stored securely by the server. The device displays necessary information only upon the user's request, and hides other data. Users have the right to view, modify, or delete their own data within the platform, and the server updates the information in the database accordingly.
[0048] Specific examples
[0049] Example 1: First conversation
[0050] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0051] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0052] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0053] Example 2: Continuation Conversation
[0054] A user types, "I've been reading mystery novels lately."
[0055] The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[0056] The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?"
[0057] The terminal displays this message to the user.
[0058] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[0062] Step 2:
[0063] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[0064] Step 3:
[0065] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[0066] Step 4:
[0067] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[0068] Step 5:
[0069] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[0070] Step 6:
[0071] The device displays the generated message to the user, who can then respond with text or voice input.
[0072] Step 7:
[0073] The device sends the user's input (text or voice) to the server.
[0074] Step 8:
[0075] The server analyzes the received input using an emotion analysis module to identify the user's emotions, such as "joy," "interest," and "sadness."
[0076] Step 9:
[0077] The server generates an empathetic response based on the extracted emotions. For example, if a user types, "I've been reading mystery novels lately," it generates a response like, "That sounds interesting. Which mystery novels did you enjoy the most?"
[0078] Step 10:
[0079] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[0080] Step 11:
[0081] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[0082] Step 12:
[0083] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[0084] Step 13:
[0085] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[0086] Step 14:
[0087] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] In today's world, there is a need to provide users who feel lonely with an effective and empathetic dialogue experience. Conventional systems have struggled to properly recognize a user's emotions and generate personalized responses. Furthermore, they lacked a means to continuously personalize the dialogue content while safely managing user information. The present invention aims to solve these problems and provide users with human-like empathy.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, terminal means for a user to access the platform and input basic information, means for transmitting the input information to the server and having the server create a profile, means for the terminal to display the generated avatar to the user and request confirmation, means for the terminal to notify the server of a user's dialogue start operation, means for transmitting user responses during the conversation from the terminal to the server, and means for generating an appropriate empathetic response based on the emotion analysis results. This enables the generation of an appropriate response according to the user's emotions and provides a personalized empathetic dialogue experience.
[0093] "Basic information" refers to information such as name, age, hobbies and interests that users enter into the system.
[0094] "Image generation artificial intelligence" is an artificial intelligence technology for generating avatars based on basic information about users.
[0095] An "avatar" is a digital representation with a character that is generated based on basic information about the user.
[0096] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates dialogue content based on user input.
[0097] "Emotion analysis" is a technology that analyzes emotions from user input and extracts meaning.
[0098] An "empathetic response" is a response that is generated in a way that is in tune with the user's emotions, based on analyzed emotional information.
[0099] The "database" is an information system that stores the generated conversation and emotion information and uses it to update the content of future conversations.
[0100] "Encryption" is a data protection method used to safely store users' personal information.
[0101] The present invention provides a system for providing personalized empathetic interaction to users who feel lonely. The system is implemented according to the following steps.
[0102] Initial Setup and Avatar Creation
[0103] Users access the platform via a web browser or a dedicated app. After accessing the platform, users enter basic information (name, age, hobbies and interests) into the displayed input form. This information is then sent from the device to the server.
[0104] The server creates a user profile based on the received basic information. Using this profile information, it activates an image generation AI (e.g., Stable Diffusion) to generate an avatar corresponding to the user. The generated avatar is sent to the terminal and displayed to the user. The user receives a display to confirm this avatar.
[0105] Conversation initiation and emotional understanding
[0106] When the user clicks the "Start" button, the operation is notified from the terminal to the server.
[0107] The server launches a conversation-generating AI (e.g., GPT-4 (registered trademark)) to generate an initial dialogue message based on the user's basic information. For example, a message like "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" is generated. The device displays the message to the user. When the user responds with text or voice, that information is sent from the device to the server.
[0108] The server analyzes the input information using an emotion analysis module (e.g., IBM Watson (registered trademark) Natural Language Understanding). Through this analysis, the server identifies the user's emotions (e.g., interest, joy). Based on the analysis results, the server generates an appropriate empathetic response. For example, if the user responds, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. What are your favorite mystery novels?" This response is displayed to the user via their terminal.
[0109] Continuous personalization
[0110] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This generates basic data to further personalize the content of the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information is stored in the database and reflected in future conversations. In the next conversation, a personalized message such as, "Regarding the painting we talked about last time, have you painted any recent works?" is generated.
[0111] Safety and Privacy Controls
[0112] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users have the right to view, correct, or delete their own data within the platform, and the server updates the information in the database accordingly.
[0113] Examples and prompts
[0114] Example 1: First conversation
[0115] 1. The user enters "Yamada Taro, 30 years old, hobby is reading."
[0116] 2. The server creates a profile for "Yamada Taro" and activates an image-generating AI (e.g., Stable Diffusion) to generate an avatar.
[0117] 3. The device displays the avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0118] Example prompt sentence:
[0119] Prompt for generative AI models (image-generating AI):
[0120] "Please create an avatar for a 30-year-old Japanese man whose hobby is reading, with a kind and intelligent face."
[0121] Prompts for generative AI models (generative conversational AI):
[0122] "The user's name is Taro Yamada, and his hobby is reading. As an initial greeting to him, generate a message that says, 'Hello, Taro Yamada. I heard you like reading. What book have you read recently?'"
[0123] Example 2: Continuation Conversation
[0124] 1. The user types, "I've been reading mystery novels lately."
[0125] 2. The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[0126] 3. The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?" and displays it on the device.
[0127] Example prompt sentence:
[0128] Prompt to the generative AI model (sentiment analysis module):
[0129] "Text: I recently read a mystery novel. Please analyze my emotions and tell me what would be an appropriate response."
[0130] Prompts for generative AI models (generative conversational AI):
[0131] "A user types, 'I've been reading mystery novels lately.' Generate an empathetic response like, 'That sounds interesting. Which mystery novels did you enjoy the most?'"
[0132] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0134] Step 1:
[0135] A user accesses the platform and enters basic information (name, age, hobbies and interests). The basic information entered by the user is sent from the device to the server. The input data includes the user's name (e.g., Yamada Taro), age (e.g., 30 years old), hobbies (e.g., reading), etc. This data is used to create the user's profile.
[0136] Input: User's basic information (name, age, hobbies and interests)
[0137] Output: Basic information sent to the server
[0138] Step 2:
[0139] The server creates a user profile based on the received basic information. This includes storing the basic information in a database. It then invokes an image-generating AI (e.g., Stable Diffusion) to generate an avatar for the user. The generated avatar has characteristics based on the basic information, such as an intelligent and kind appearance. The generated avatar is sent from the server to the device and displayed to the user. It is displayed along with the message, "Are you sure you want this avatar?"
[0140] Input: Basic information sent to the server
[0141] Output: An avatar image and a confirmation message that will be displayed to the user.
[0142] Step 3:
[0143] When the user clicks the "Start" button, the device notifies the server of the operation. The server receives this notification and launches a conversation-generating AI (e.g., GPT-4). The server generates an initial greeting and question based on the user's basic information. For example, it generates a message like, "Hello, Yamada-san. I heard you like reading. What book have you read recently?" This message is sent from the server to the device and displayed to the user.
[0144] Input: User interaction start operation
[0145] Output: Initial greeting message and question
[0146] Step 4:
[0147] When the user responds with text or voice, the input is sent from the device to the server. The server analyzes the input using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). The analysis identifies the user's emotion (e.g., interest, joy). The extracted emotion information is used in the next step.
[0148] Input: User response (text or voice)
[0149] Output: Parsed emotion data
[0150] Step 5:
[0151] The server generates an appropriate empathetic response based on the results of the emotion analysis. It uses conversation-generating AI to create an empathetic message. For example, a response might be generated such as, "That sounds interesting. What are your favorite mystery novels?" This response is sent from the server to the device and displayed to the user.
[0152] Input: Parsed emotion data
[0153] Output: Empathetic response message
[0154] Step 6:
[0155] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This stored data is used to personalize the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information will be reflected in the next conversation. In the next conversation, a personalized message will be generated, such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[0156] Input: Conversation history and emotional information
[0157] Output: Database information used for future personalization
[0158] Step 7:
[0159] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users can view, modify, and delete their own data within the platform, and the server updates the information in the database accordingly.
[0160] Input: User's personal information
[0161] Output: Encrypted database information and the ability to view, modify, and delete data according to user requests
[0162] (Application example 1)
[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0164] In modern society, the number of users who feel lonely is increasing. To improve customer experience, particularly in brick-and-mortar stores, personalized, empathetic dialogue is needed, rather than simply providing guidance. However, current in-store guidance systems and guide robots have difficulty responding flexibly based on the customer's basic information and emotions. Therefore, the present invention aims to provide a system that reduces customers' feelings of loneliness and improves customer experience by providing personalized, empathetic dialogue.
[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0166] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for providing personalized guidance through continuous dialogue based on the user's basic information, means for implementing the system in a store using a robot, and means for encrypting the user's personal information and securely storing the avatar and its data. This makes it possible to provide personalized empathetic dialogue in a store based on the user's basic information, hobbies, and interests.
[0167] "Basic information" refers to the initial data needed to individualize and personalize the interaction, such as the user's name, age, hobbies, and interests.
[0168] "Image generation artificial intelligence" is an artificial intelligence technology for generating an avatar corresponding to a user based on basic information about the user.
[0169] An "avatar" is a digital person or character with a virtual appearance and personality that is generated based on basic information about the user.
[0170] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates appropriate dialogue content based on input information from the user.
[0171] "Emotion analysis" is a technology for extracting emotions from user input and determining emotions such as interest, joy, sadness, etc.
[0172] An "empathetic response" is a response that is generated based on the results of analyzing the user's emotions and is given in a manner that empathizes with the user's emotions.
[0173] The "database" is an information management system that stores the generated conversation and emotion information and updates or references it as needed.
[0174] "Encryption" is the technology of converting data into cryptographic code to protect users' personal information.
[0175] "Individualized guidance" refers to guidance information tailored to the individual needs and interests of a user, provided based on the user's basic information and past interaction history.
[0176] A "robot" is a mechanical device that exists physically and has a guidance function to operate the above system in the real world.
[0177] To implement the present invention, the following system configuration and processing procedures are basically required.
[0178] A user first accesses the system and enters their basic information, including name, age, hobbies, and interests. The system then creates a user profile and uses image generation AI to generate an avatar for the user. The terminal then displays the generated avatar to the user and asks for their confirmation.
[0179] When the user clicks the "Start" button to begin a conversation with the avatar, the device notifies the server of this action. The server then activates a conversation-generating AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" may be generated.
[0180] When the user responds with text or voice, the device sends the input to the server. The server analyzes the input using an emotion analysis module to extract the user's emotion. An empathetic response is generated based on the analyzed emotion. For example, if the user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via the device.
[0181] This system realizes continuous personalized dialogue by storing the content of the dialogue with the user and emotional information in a database. The next time the dialogue is held, the content is updated based on the stored data, providing a more personalized response. For example, if the user mentioned a new hobby (e.g., painting) in the previous dialogue, the next dialogue could continue with a question such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[0182] The system encrypts users' personal information and stores avatars and related data securely, protecting users' privacy. Users can also view, modify, and delete their own data within the platform.
[0183] As a concrete example, the following system usage scenarios can be considered:
[0184] Example 1: When a guide robot approaches a customer in a store
[0185] "Hello, Mr. / Ms. XX. Thank you for coming. I heard you enjoy reading. What book have you read recently?"
[0186] An example of a prompt is:
[0187] User: Yamada Taro. Age: 30. Hobbies: reading and watching movies.
[0188] What mystery novel have you read recently?”
[0189] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0191] Step 1:
[0192] A user accesses the system and enters their basic information. This basic information includes name, age, hobbies, interests, etc. The entered basic information is sent from the terminal to the server. Input: User's basic information. Output: Basic information saved on the server.
[0193] Step 2:
[0194] The server creates a user profile based on the received basic information. An image generation AI uses this user profile to generate an avatar. The generated avatar is sent to the device. Input: Basic user information. Output: Generated avatar.
[0195] Step 3:
[0196] The device displays the generated avatar to the user and asks for confirmation. Input: Generated avatar. Output: Displayed avatar and user confirmation.
[0197] Step 4:
[0198] When the user clicks the "Start" button to start interacting with the avatar, the terminal notifies the server of this operation. Input: User operation. Output: Notification to the server.
[0199] Step 5:
[0200] The server starts a conversation generation AI to generate an initial greeting and question based on the user's basic information. The generated message is sent to the terminal and displayed to the user. Input: User's basic information. Output: Generated initial message.
[0201] Step 6:
[0202] The user responds with text or voice. This input is sent from the device to the server. Input: User response. Output: Response sent to the server.
[0203] Step 7:
[0204] The server analyzes the input content with an emotion analysis module to extract the user's emotion. The analysis result is used as input for generating an empathetic response. Input: User's response. Output: Analyzed emotion.
[0205] Step 8:
[0206] The server generates an empathetic response based on the emotion analysis results and sends the response to the device. The device displays the generated response to the user. Input: Analyzed emotion. Output: Generated empathetic response.
[0207] Step 9:
[0208] The server saves the generated conversation and emotion information in a database. The next time the conversation is held, the content of the conversation will be updated based on this information. Input: Generated conversation and emotion information. Output: Information saved in the database.
[0209] Step 10:
[0210] The server encrypts all data and stores it securely. Users can view, modify, or delete their own data at any time. Input: Information stored in the database. Output: Encrypted data.
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely, and by combining an emotion engine, recognizes the user's emotions and generates responses accordingly. The program and processing flow of this system are described in detail below.
[0213] Initial Setup and Avatar Creation
[0214] First, the user accesses the platform and enters basic information (such as name, age, hobbies and interests). The device then sends this information to the server. The server then creates a user profile based on the received basic information and calls on image-generating AI to generate an avatar. The avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[0215] Conversation initiation and emotion recognition
[0216] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, it might generate a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" In response, the user inputs text or voice, and the device sends the input to the server.
[0217] Emotion recognition by emotion engine
[0218] The server analyzes the received input content using an emotion engine to identify the user's emotion (e.g., interest, joy, sadness). The emotion engine extracts emotional signals from various inputs such as text, voice, and facial expressions to determine the user's emotional state. For example, if a user inputs, "I've been reading mystery novels recently," the emotion engine will identify "interest" and "joy."
[0219] Generating empathetic responses and interacting
[0220] Based on this analyzed emotional information, the server generates an empathetic response, such as, "That sounds interesting. Which mystery novel did you particularly enjoy?" Furthermore, the avatar's facial expression and attitude change according to this emotional information, visually demonstrating empathy.
[0221] The generated response message is displayed to the user via the terminal, and this dialogue cycle is repeated.Furthermore, the server stores the generated conversation and emotion information in a database and manages the conversation history.
[0222] Continuous personalization
[0223] The server optimizes the content of the next conversation based on the history in the database. It learns changes in the user's hobbies and interests and reflects these in the content of the conversation. For example, if the user talks about a new hobby (e.g., painting), the server continues the conversation by asking, "Regarding the painting we talked about last time, have you painted any works recently?"
[0224] Safety and Privacy Controls
[0225] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon the user's request, and hides other data. Users can view their own data within the platform and modify or delete it as necessary. The server updates the information in the database in response to this request, protecting privacy.
[0226] Specific examples
[0227] Example 1: First conversation
[0228] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0229] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0230] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0231] Example 2: Continuation Conversation
[0232] A user types, "I've been reading mystery novels lately."
[0233] The server analyzes the input using an emotion engine and determines it as "interest."
[0234] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0235] The terminal displays this message to the user.
[0236] In this way, the system provides a personalized and empathetic interaction experience for the user, reducing feelings of loneliness and leveraging the emotion engine to further address the user's emotional state.
[0237] The processing flow will be explained below.
[0238] Step 1:
[0239] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[0240] Step 2:
[0241] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[0242] Step 3:
[0243] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[0244] Step 4:
[0245] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[0246] Step 5:
[0247] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[0248] Step 6:
[0249] The device displays the generated message to the user, who can then respond with text or voice input.
[0250] Step 7:
[0251] The device sends the user's input (text or voice) to the server.
[0252] Step 8:
[0253] The server analyzes the received input using an emotion engine to identify the user's emotions, such as "joy," "interest," and "sadness."
[0254] Step 9:
[0255] The server then sends commands to the device to change the avatar's facial expression and attitude based on the extracted emotional information. For example, if the user inputs "I've been reading mystery novels recently," the avatar will display expressions and attitudes that express "interest" or "joy."
[0256] Step 10:
[0257] The server generates an empathetic response based on the emotional information, for example, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0258] Step 11:
[0259] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[0260] Step 12:
[0261] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[0262] Step 13:
[0263] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[0264] Step 14:
[0265] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[0266] Step 15:
[0267] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[0268] Specific examples
[0269] Example 1: First conversation
[0270] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0271] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0272] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0273] Example 2: Continuation Conversation
[0274] A user types, "I've been reading mystery novels lately."
[0275] The server analyzes the input using an emotion engine and determines it as "interest."
[0276] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0277] The device displays a response message, and the avatar's facial expression and attitude change to reflect interest.
[0278] Example 2
[0279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0280] In modern society, the number of people feeling lonely is increasing, and there is a need for systems that provide empathetic dialogue tailored to individual needs. However, conventional dialogue systems have difficulty fully analyzing users' emotions and providing empathetic responses. They also fall short in terms of protecting the safety and privacy of users' personal information. A system that solves these problems and reduces feelings of loneliness is needed.
[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0282] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next dialogue based on the stored information, means for encrypting the user's personal information and safely storing the data, means for changing the avatar's facial expression and attitude based on the emotional information, means for displaying necessary information and hiding other data in response to a user request, and means for the user to view, modify, and delete data. This makes it possible to provide an empathetic dialogue that is responsive to the user's emotions, while protecting the safety and privacy of the personal information.
[0283] "User" refers to an individual who interacts with the system.
[0284] "Basic Information" refers to personal information entered by the user, such as name, age, hobbies and interests.
[0285] A "terminal" is a device that is operated by a user to input and display information.
[0286] "Server" refers to a central computer system that receives and processes information sent by users.
[0287] "Image generation artificial intelligence" refers to artificial intelligence technology for generating avatars based on basic user information.
[0288] An "avatar" is a virtual persona or character created based on a user's basic information.
[0289] "Conversational generative artificial intelligence" refers to artificial intelligence that has the technology to generate dialogue content based on user input.
[0290] An "emotion engine" refers to algorithms and technologies for analyzing and identifying emotions from user input.
[0291] "Emotion information" refers to data on the user's emotional state analyzed by the emotion engine.
[0292] "Empathetic response" refers to a dialogue that is based on analyzed emotional information and is sensitive to the user's emotions.
[0293] "Database" refers to a data structure for storing and managing generated conversation and emotion information.
[0294] "Encryption" refers to the technology of converting digital data into a form that cannot be read by third parties.
[0295] "Safe storage" refers to measures to protect users' personal information and data from unauthorized access and information leaks.
[0296] "Changing facial expressions and attitudes" refers to changing the visual response of the avatar displayed to the user based on emotional information.
[0297] "Displaying information on demand" refers to displaying only the necessary information based on the user's instructions.
[0298] "Viewing, modifying, and deleting data" refers to the actions that users can take to view, modify, or delete their own data through the platform.
[0299] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely. This system has the following configuration to recognize the user's emotions by combining emotion engines and generate responses according to those emotions.
[0300] Initial Setup and Avatar Creation
[0301] First, a user accesses the platform and enters basic information such as name, age, hobbies, and interests. The device then sends this information to the server. The server creates a user profile based on the received information and calls on image-generating AI (e.g., DeepArt) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input information.
[0302] The generated avatar is displayed to the user via the terminal and they are asked to confirm. Through this process, the user can obtain an avatar that suits their preferences.
[0303] Conversation initiation and emotion recognition
[0304] To begin a conversation with an avatar, a user clicks the "Start" button. This action causes the device to notify the server. The server then launches a conversation-generating AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[0305] For example, the server might generate a message like, "Hello, [username]. I heard you like reading. What book have you read recently?" The user can respond with text or voice, and the device will send that input to the server.
[0306] Emotion recognition by emotion engine
[0307] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer), which extracts emotional signals from inputs such as text, voice, and facial expressions to identify the user's emotions (interest, joy, sadness, etc.).
[0308] For example, if a user types, "I recently read a mystery novel," the emotion engine identifies emotions like "interest" and "joy," and based on this, the server generates a more empathetic response.
[0309] Generating empathetic responses and interacting
[0310] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds interesting. Which mystery novel did you particularly enjoy?" The avatar's facial expressions and attitude also change according to this emotional information, visually demonstrating empathy.
[0311] The generated responses are displayed to the user via the device, and this dialogue cycle is repeated. At the same time, the server stores the generated conversation and emotional information in a database and manages the conversation history. This history is used to optimize the content of the next dialogue.
[0312] Continuous personalization
[0313] The server optimizes the content of the next conversation based on the saved history data. This allows the server to learn about changes in the user's hobbies and interests and reflect that information in the conversation. For example, if the user takes up a new hobby (e.g., painting), the conversation can continue with a prompt such as, "Regarding the painting we talked about last time, have you painted any recent works?"
[0314] Safety and Privacy Controls
[0315] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon user request, hiding other data. Furthermore, users can view their own data within the platform and modify or delete it as needed. The server updates the information in the database in response to these requests, protecting privacy.
[0316] Specific examples
[0317] Example 1: First conversation
[0318] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0319] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0320] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0321] Example 2: Continuation Conversation
[0322] A user types, "I've been reading mystery novels lately."
[0323] The server analyzes the input using an emotion engine and determines it as "interest."
[0324] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0325] The terminal displays this message to the user.
[0326] In this way, the system of the present invention can provide a personalized and empathetic interaction experience for the user, reducing feelings of loneliness. Furthermore, by utilizing the emotion engine, the system can generate appropriate responses to the user's emotional state, improving the quality of the interaction.
[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0328] Step 1: Enter basic information
[0329] Users access the platform and enter basic information such as their name, age, hobbies and interests.
[0330] Input: Basic information such as name, age, hobbies and interests
[0331] Output: Basic information data
[0332] Step 2: Submit basic information
[0333] The terminal transmits the input basic information to the server.
[0334] Input: Basic information data
[0335] Output: Basic information data sent to the server
[0336] Step 3: Create your profile and create your avatar
[0337] The server creates a user profile based on the received basic information and calls an image-generating artificial intelligence (e.g., DeepArt) to generate an avatar.
[0338] Input: Basic information data sent to the server
[0339] Output: Generated avatar data
[0340] Step 4: View and check your avatar
[0341] The terminal displays the generated avatar to the user and asks for confirmation.
[0342] Input: Generated avatar data
[0343] Output: The avatar displayed to the user
[0344] Step 5: Initiating an initial conversation
[0345] The user clicks the "Get Started" button.
[0346] The terminal notifies the server of this operation.
[0347] Input: User clicks the "Get Started" button
[0348] Output: Notification to the server
[0349] Step 6: Generate initial greetings and questions
[0350] The server launches a conversational AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[0351] Input: Notification to the server, basic information data of the user
[0352] Output: Initial greeting and question message
[0353] Step 7: User response input
[0354] The user enters a response by text or voice.
[0355] The terminal transmits the input response content to the server.
[0356] Input: User response
[0357] Output: Response sent to the server
[0358] Step 8: Emotion Recognition
[0359] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the emotion.
[0360] Input: Response sent to the server
[0361] Output: Identified emotion information
[0362] Step 9: Generate an empathetic response
[0363] The server generates an empathetic response based on the analyzed emotional information, such as "That sounds interesting. Which mystery novel did you particularly enjoy?", and also changes the avatar's facial expression and attitude according to this emotional information.
[0364] Input: Identified emotion information
[0365] Output: Empathetic response message, modified avatar data
[0366] Step 10: View the response
[0367] The generated response message is displayed to the user through the terminal, along with the avatar's facial expression and attitude.
[0368] Input: Empathetic response message, modified avatar data
[0369] Output: Response message displayed to the user and the avatar's facial expression and attitude
[0370] Step 11: Saving conversation history
[0371] The server stores the generated conversation and emotion information in a database.
[0372] Input: Generated conversation data, identified emotion information
[0373] Output: Conversation data and emotion information stored in a database
[0374] Step 12: Optimize the content of your next conversation
[0375] The server optimizes the content of the next conversation based on the stored historical data.
[0376] Input: Conversation data and emotion information stored in a database
[0377] Output: Optimized next dialogue
[0378] Step 13: Managing your personal information
[0379] The server encrypts and securely stores users' personal information.
[0380] The device displays necessary information only when requested by the user, and hides other data. The user can also view, modify, and delete data.
[0381] Input: User's personal information, User's request
[0382] Output: Encrypted personal information, securely stored data, data displayed, modified, or deleted upon user request
[0383] Through these steps, it is possible to provide users with personalized, empathetic interactions and reduce their feelings of loneliness. Furthermore, by utilizing the emotion engine, it is possible to generate appropriate responses according to the user's emotional state, improving the quality of the interactions.
[0384] (Application example 2)
[0385] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0386] Conventional food delivery applications lacked a way to address the feelings of loneliness felt by users when ordering food. Furthermore, they were unable to provide empathetic dialogue that responded to the user's emotional state, making it difficult to improve the quality of the user experience.
[0387] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0388] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the next conversation content based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, means for providing empathetic dialogue if the user feels lonely based on the input content when ordering a meal, and means for analyzing the user's emotions when ordering a meal using an emotion engine and suggesting a response based on the emotions. This makes it possible to provide psychological support through empathetic dialogue even when the user feels lonely, thereby improving the quality of the user experience.
[0389] "Basic information" is personal information such as the user's name, age, hobbies and interests.
[0390] "Image generation artificial intelligence" is an artificial intelligence technology that generates images based on input data.
[0391] An "avatar" is a virtual character created to represent a user.
[0392] "Conversational artificial intelligence" is an artificial intelligence technology that generates natural dialogue based on user input.
[0393] An "emotion engine" is a technology for analyzing emotions from user input and behavior.
[0394] An "empathetic response" is a response that understands the user's emotions and is generated in a way that is sympathetic to those emotions.
[0395] The "database" is an information storage system for storing users' conversation history and emotional information.
[0396] "Encryption" is a technology that converts data to safely store users' personal information and prevents external access.
[0397] "Food delivery" is a service that allows users to order meals online and have them delivered to a specified location.
[0398] "Loneliness" is a psychological state in which one feels disconnected from others and isolated.
[0399] "Psychological support" is assistance to provide users with psychological stability and a sense of security.
[0400] This invention is a system for realizing a food delivery application that provides empathetic dialogue when a user feels lonely when ordering a meal. This system operates using a smartphone and a server.
[0401] First, a user accesses the application using their smartphone and enters basic information such as name, age, hobbies and interests, which is then sent from the device to a server, which then creates a user profile.
[0402] Next, the server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input. The generated avatar is displayed on the device and the user is asked to confirm it.
[0403] When a user starts interacting with an avatar, the server activates a conversation-generating AI (e.g., GPT-3 (registered trademark)) to generate an initial greeting and question. Based on the user's basic information, the server generates empathetic messages such as "Hello, how busy have you been?" or "Do you have any particular favorite dishes?" The user inputs the information as text or voice, which is then sent from the device to the server.
[0404] The server analyzes the user's input using an emotion engine (e.g., IBM Watson or Microsoft® Azure® Emotion API) to identify the user's emotion. For example, if a user inputs, "I often eat alone, and I feel a little lonely," the emotion engine will identify the emotion as "loneliness."
[0405] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds lonely. What kind of person would be good to be with at a time like this?" The generated response is then displayed to the user via the device.
[0406] This cycle repeats, and the server stores the generated conversation and emotional information in a database, optimizing the content of the next conversation based on the stored information and adapting to changes in the user's hobbies and interests.
[0407] In addition, users' personal information is encrypted and securely stored by the server. This allows users to view, modify, and delete their data within the platform. The device will only display necessary information upon user request, and will hide other data.
[0408] As a concrete example, the following cases can be considered:
[0409] Example 1: First conversation
[0410] The user types, "Hello, I'm Taro Tanaka, I'm 25 years old, and my hobby is cooking."
[0411] The server generates the message "Hello, Tanaka-san. I heard you like cooking. What dish have you made recently?"
[0412] Example 2: Disconnected conversations
[0413] The user types, "I often eat alone and feel a little lonely."
[0414] The server generates an empathetic response: "That sounds lonely. Who do you think would be good to have around at a time like this?"
[0415] Prompt Sentence Examples
[0416] User: "Hi, I'm very tired."
[0417] AI: "I see, that's tough. What was particularly tiring for you today?"
[0418] In this way, the system of the present invention provides users with an empathetic interactive experience to reduce feelings of loneliness and realize psychological support.
[0419] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0420] Step 1:
[0421] The user uses a smartphone to enter basic information (name, age, hobbies, interests, etc.).
[0422] Input: Name, age, hobbies and interests
[0423] Output: Basic information set (user information completed)
[0424] What happens: A user enters the required information into a form and clicks the submit button.
[0425] Step 2:
[0426] The terminal sends the user's basic information to the server.
[0427] Input: Basic information
[0428] Output: Send data to the server
[0429] Operation: The device encrypts the basic information entered and sends it to the server.
[0430] Step 3:
[0431] The server uses image generation artificial intelligence to generate an avatar based on basic information.
[0432] Input: Encrypted basic information
[0433] Output: The generated avatar
[0434] How it works: The server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar that reflects the user's characteristics.
[0435] Step 4:
[0436] The terminal displays the generated avatar to the user and asks for confirmation.
[0437] Input: Generated avatar
[0438] Output: Display avatar and confirm user
[0439] What it does: The device displays the avatar on the screen and prompts the user to confirm the avatar.
[0440] Step 5:
[0441] The user clicks the "Start" button to begin interacting with the avatar.
[0442] Input: Start Operation
[0443] Output: Trigger to start a conversation
[0444] Action: The user clicks the "Get Started" button on the application screen.
[0445] Step 6:
[0446] The server activates a conversation-generating artificial intelligence (AI) to generate initial greetings and questions based on the user's basic information.
[0447] Input: Basic information, trigger to start a conversation
[0448] Output: Initial greetings and questions
[0449] How it works: The server uses a generative conversational AI (e.g., GPT-3) to generate messages such as "Hello, how have you been?"
[0450] Step 7:
[0451] The user inputs text or voice, and the device sends the input to the server.
[0452] Input: User text or voice input
[0453] Output: Sending user input to the server
[0454] Operation: The device sends the user's text or voice data to the server.
[0455] Step 8:
[0456] The server analyzes the user's input using an emotion engine and identifies the emotion.
[0457] Input: User text or voice input
[0458] Output: Identified emotion information
[0459] How it works: The server uses an emotion engine (e.g. IBM Watson or Microsoft Azure Emotion API) to analyze emotions from the input.
[0460] Step 9:
[0461] Based on the analyzed emotional information, the server generates an empathetic response.
[0462] Input: Emotion information
[0463] Output: Empathetic response message
[0464] How it works: The server generates an empathetic message based on the emotional information, for example, a response like "That sounds lonely."
[0465] Step 10:
[0466] The generated empathetic response is displayed to the user via the terminal.
[0467] Input: Empathetic response message
[0468] Output: The response message that is displayed to the user
[0469] Action: The terminal displays the response message on the screen.
[0470] Step 11:
[0471] The server stores the generated conversation and emotion information in a database.
[0472] Input: Conversation content, emotional information
[0473] Output: Save to database
[0474] How it works: The server records conversation history and emotion information in a database.
[0475] Step 12:
[0476] The conversation will be updated based on the saved information the next time you interact, and will also respond to changes in the user's hobbies and interests.
[0477] Input: saved conversation history, emotional information
[0478] Output: Optimized dialogue
[0479] How it works: The server retrieves information from the database and personalizes your next interaction.
[0480] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0481] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0482] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0483] [Second embodiment]
[0484] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0485] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0486] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0487] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0488] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0489] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0490] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0491] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0492] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0493] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0494] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0495] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0496] The present invention is a system that provides personalized, empathetic dialogue to users who feel lonely. This system uses artificial intelligence to prompt users to enter basic information and engage in dialogue through a generated avatar. It also analyzes the user's emotions and generates responses based on those emotions. The program for this system is described in detail below.
[0497] Initial Setup and Avatar Creation
[0498] First, a user accesses the platform and enters basic information (name, age, hobbies and interests). This basic information is sent from the device to the server, which then creates a user profile based on it. Next, the server uses image generation artificial intelligence to generate an avatar for the user. This avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[0499] Conversation initiation and emotional understanding
[0500] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" can be generated. In response, the user inputs text or voice, and the device sends the input to the server.
[0501] The server analyzes this input using an emotion analysis module to extract the user's emotions (e.g., interest, joy). It then generates an empathetic response based on the analyzed emotions. For example, if a user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via their device.
[0502] Continuous personalization
[0503] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This allows the server to have basic data to better personalize the content of the next conversation. For example, if the user talks about a new hobby (e.g., painting), this information will be reflected in future conversations. The conversation can continue in the form of, "Regarding the painting we talked about last time, have you painted any works recently?"
[0504] Safety and Privacy Controls
[0505] Users' personal information is encrypted and stored securely by the server. The device displays necessary information only upon the user's request, and hides other data. Users have the right to view, modify, or delete their own data within the platform, and the server updates the information in the database accordingly.
[0506] Specific examples
[0507] Example 1: First conversation
[0508] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0509] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0510] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0511] Example 2: Continuation Conversation
[0512] A user types, "I've been reading mystery novels lately."
[0513] The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[0514] The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?"
[0515] The terminal displays this message to the user.
[0516] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0517] The processing flow will be explained below.
[0518] Step 1:
[0519] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[0520] Step 2:
[0521] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[0522] Step 3:
[0523] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[0524] Step 4:
[0525] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[0526] Step 5:
[0527] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[0528] Step 6:
[0529] The device displays the generated message to the user, who can then respond with text or voice input.
[0530] Step 7:
[0531] The device sends the user's input (text or voice) to the server.
[0532] Step 8:
[0533] The server analyzes the received input using an emotion analysis module to identify the user's emotions, such as "joy," "interest," and "sadness."
[0534] Step 9:
[0535] The server generates an empathetic response based on the extracted emotions. For example, if a user types, "I've been reading mystery novels lately," it generates a response like, "That sounds interesting. Which mystery novels did you enjoy the most?"
[0536] Step 10:
[0537] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[0538] Step 11:
[0539] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[0540] Step 12:
[0541] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[0542] Step 13:
[0543] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[0544] Step 14:
[0545] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[0546] Example 1
[0547] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0548] In today's world, there is a need to provide users who feel lonely with an effective and empathetic dialogue experience. Conventional systems have struggled to properly recognize a user's emotions and generate personalized responses. Furthermore, they lacked a means to continuously personalize the dialogue content while safely managing user information. The present invention aims to solve these problems and provide users with human-like empathy.
[0549] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0550] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, terminal means for a user to access the platform and input basic information, means for transmitting the input information to the server and having the server create a profile, means for the terminal to display the generated avatar to the user and request confirmation, means for the terminal to notify the server of a user's dialogue start operation, means for transmitting user responses during the conversation from the terminal to the server, and means for generating an appropriate empathetic response based on the emotion analysis results. This enables the generation of an appropriate response according to the user's emotions and provides a personalized empathetic dialogue experience.
[0551] "Basic information" refers to information such as name, age, hobbies and interests that users enter into the system.
[0552] "Image generation artificial intelligence" is an artificial intelligence technology for generating avatars based on basic information about users.
[0553] An "avatar" is a digital representation with a character that is generated based on basic information about the user.
[0554] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates dialogue content based on user input.
[0555] "Emotion analysis" is a technology that analyzes emotions from user input and extracts meaning.
[0556] An "empathetic response" is a response that is generated in a way that is in tune with the user's emotions, based on analyzed emotional information.
[0557] The "database" is an information system that stores the generated conversation and emotion information and uses it to update the content of future conversations.
[0558] "Encryption" is a data protection method used to safely store users' personal information.
[0559] The present invention provides a system for providing personalized empathetic interaction to users who feel lonely. The system is implemented according to the following steps.
[0560] Initial Setup and Avatar Creation
[0561] Users access the platform via a web browser or a dedicated app. After accessing the platform, users enter basic information (name, age, hobbies and interests) into the displayed input form. This information is then sent from the device to the server.
[0562] The server creates a user profile based on the received basic information. Using this profile information, it activates an image generation AI (e.g., Stable Diffusion) to generate an avatar corresponding to the user. The generated avatar is sent to the terminal and displayed to the user. The user receives a display to confirm this avatar.
[0563] Conversation initiation and emotional understanding
[0564] When the user clicks the "Start" button, the operation is notified from the terminal to the server.
[0565] The server launches a conversational AI (e.g., GPT-4) to generate an initial dialogue message based on the user's basic information. For example, a message like "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" is generated. The device displays the message to the user. When the user responds with text or voice, that information is sent from the device to the server.
[0566] The server analyzes the input information using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). Through this analysis, the server identifies the user's emotions (e.g., interest, joy). Based on the analysis results, the server generates an appropriate empathetic response. For example, if the user responds, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. What are your favorite mystery novels?" This response is displayed to the user via their device.
[0567] Continuous personalization
[0568] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This generates basic data to further personalize the content of the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information is stored in the database and reflected in future conversations. In the next conversation, a personalized message such as, "Regarding the painting we talked about last time, have you painted any recent works?" is generated.
[0569] Safety and Privacy Controls
[0570] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users have the right to view, correct, or delete their own data within the platform, and the server updates the information in the database accordingly.
[0571] Examples and prompts
[0572] Example 1: First conversation
[0573] 1. The user enters "Yamada Taro, 30 years old, hobby is reading."
[0574] 2. The server creates a profile for "Yamada Taro" and activates an image-generating AI (e.g., Stable Diffusion) to generate an avatar.
[0575] 3. The device displays the avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0576] Example prompt sentence:
[0577] Prompt for generative AI models (image-generating AI):
[0578] "Please create an avatar for a 30-year-old Japanese man whose hobby is reading, with a kind and intelligent face."
[0579] Prompts for generative AI models (generative conversational AI):
[0580] "The user's name is Taro Yamada, and his hobby is reading. As an initial greeting to him, generate a message that says, 'Hello, Taro Yamada. I heard you like reading. What book have you read recently?'"
[0581] Example 2: Continuation Conversation
[0582] 1. The user types, "I've been reading mystery novels lately."
[0583] 2. The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[0584] 3. The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?" and displays it on the device.
[0585] Example prompt sentence:
[0586] Prompt to the generative AI model (sentiment analysis module):
[0587] "Text: I recently read a mystery novel. Please analyze my emotions and tell me what would be an appropriate response."
[0588] Prompts for generative AI models (generative conversational AI):
[0589] "A user types, 'I've been reading mystery novels lately.' Generate an empathetic response like, 'That sounds interesting. Which mystery novels did you enjoy the most?'"
[0590] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0591] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0592] Step 1:
[0593] A user accesses the platform and enters basic information (name, age, hobbies and interests). The basic information entered by the user is sent from the device to the server. The input data includes the user's name (e.g., Yamada Taro), age (e.g., 30 years old), hobbies (e.g., reading), etc. This data is used to create the user's profile.
[0594] Input: User's basic information (name, age, hobbies and interests)
[0595] Output: Basic information sent to the server
[0596] Step 2:
[0597] The server creates a user profile based on the received basic information. This includes storing the basic information in a database. It then invokes an image-generating AI (e.g., Stable Diffusion) to generate an avatar for the user. The generated avatar has characteristics based on the basic information, such as an intelligent and kind appearance. The generated avatar is sent from the server to the device and displayed to the user. It is displayed along with the message, "Are you sure you want this avatar?"
[0598] Input: Basic information sent to the server
[0599] Output: An avatar image and a confirmation message that will be displayed to the user.
[0600] Step 3:
[0601] When the user clicks the "Start" button, the device notifies the server of the operation. The server receives this notification and launches a conversation-generating AI (e.g., GPT-4). The server generates an initial greeting and question based on the user's basic information. For example, it generates a message like, "Hello, Yamada-san. I heard you like reading. What book have you read recently?" This message is sent from the server to the device and displayed to the user.
[0602] Input: User interaction start operation
[0603] Output: Initial greeting message and question
[0604] Step 4:
[0605] When the user responds with text or voice, the input is sent from the device to the server. The server analyzes the input using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). The analysis identifies the user's emotion (e.g., interest, joy). The extracted emotion information is used in the next step.
[0606] Input: User response (text or voice)
[0607] Output: Parsed emotion data
[0608] Step 5:
[0609] The server generates an appropriate empathetic response based on the results of the emotion analysis. It uses conversation-generating AI to create an empathetic message. For example, a response might be generated such as, "That sounds interesting. What are your favorite mystery novels?" This response is sent from the server to the device and displayed to the user.
[0610] Input: Parsed emotion data
[0611] Output: Empathetic response message
[0612] Step 6:
[0613] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This stored data is used to personalize the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information will be reflected in the next conversation. In the next conversation, a personalized message will be generated, such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[0614] Input: Conversation history and emotional information
[0615] Output: Database information used for future personalization
[0616] Step 7:
[0617] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users can view, modify, and delete their own data within the platform, and the server updates the information in the database accordingly.
[0618] Input: User's personal information
[0619] Output: Encrypted database information and the ability to view, modify, and delete data according to user requests
[0620] (Application example 1)
[0621] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0622] In modern society, the number of users who feel lonely is increasing. To improve customer experience, particularly in brick-and-mortar stores, personalized, empathetic dialogue is needed, rather than simply providing guidance. However, current in-store guidance systems and guide robots have difficulty responding flexibly based on the customer's basic information and emotions. Therefore, the present invention aims to provide a system that reduces customers' feelings of loneliness and improves customer experience by providing personalized, empathetic dialogue.
[0623] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0624] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for providing personalized guidance through continuous dialogue based on the user's basic information, means for implementing the system in a store using a robot, and means for encrypting the user's personal information and securely storing the avatar and its data. This makes it possible to provide personalized empathetic dialogue in a store based on the user's basic information, hobbies, and interests.
[0625] "Basic information" refers to the initial data needed to individualize and personalize the interaction, such as the user's name, age, hobbies, and interests.
[0626] "Image generation artificial intelligence" is an artificial intelligence technology for generating an avatar corresponding to a user based on basic information about the user.
[0627] An "avatar" is a digital person or character with a virtual appearance and personality that is generated based on basic information about the user.
[0628] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates appropriate dialogue content based on input information from the user.
[0629] "Emotion analysis" is a technology for extracting emotions from user input and determining emotions such as interest, joy, sadness, etc.
[0630] An "empathetic response" is a response that is generated based on the results of analyzing the user's emotions and is given in a manner that empathizes with the user's emotions.
[0631] The "database" is an information management system that stores the generated conversation and emotion information and updates or references it as needed.
[0632] "Encryption" is the technology of converting data into cryptographic code to protect users' personal information.
[0633] "Individualized guidance" refers to guidance information tailored to the individual needs and interests of a user, provided based on the user's basic information and past interaction history.
[0634] A "robot" is a mechanical device that exists physically and has a guidance function to operate the above system in the real world.
[0635] To implement the present invention, the following system configuration and processing procedures are basically required.
[0636] A user first accesses the system and enters their basic information, including name, age, hobbies, and interests. The system then creates a user profile and uses image generation AI to generate an avatar for the user. The terminal then displays the generated avatar to the user and asks for their confirmation.
[0637] When the user clicks the "Start" button to begin a conversation with the avatar, the device notifies the server of this action. The server then activates a conversation-generating AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" may be generated.
[0638] When the user responds with text or voice, the device sends the input to the server. The server analyzes the input using an emotion analysis module to extract the user's emotion. An empathetic response is generated based on the analyzed emotion. For example, if the user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via the device.
[0639] This system realizes continuous personalized dialogue by storing the content of the dialogue with the user and emotional information in a database. The next time the dialogue is held, the content is updated based on the stored data, providing a more personalized response. For example, if the user mentioned a new hobby (e.g., painting) in the previous dialogue, the next dialogue could continue with a question such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[0640] The system encrypts users' personal information and stores avatars and related data securely, protecting users' privacy. Users can also view, modify, and delete their own data within the platform.
[0641] As a concrete example, the following system usage scenarios can be considered:
[0642] Example 1: When a guide robot approaches a customer in a store
[0643] "Hello, Mr. / Ms. XX. Thank you for coming. I heard you enjoy reading. What book have you read recently?"
[0644] An example of a prompt is:
[0645] User: Yamada Taro. Age: 30. Hobbies: reading and watching movies.
[0646] What mystery novel have you read recently?”
[0647] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0648] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0649] Step 1:
[0650] A user accesses the system and enters their basic information. This basic information includes name, age, hobbies, interests, etc. The entered basic information is sent from the terminal to the server. Input: User's basic information. Output: Basic information saved on the server.
[0651] Step 2:
[0652] The server creates a user profile based on the received basic information. An image generation AI uses this user profile to generate an avatar. The generated avatar is sent to the device. Input: Basic user information. Output: Generated avatar.
[0653] Step 3:
[0654] The device displays the generated avatar to the user and asks for confirmation. Input: Generated avatar. Output: Displayed avatar and user confirmation.
[0655] Step 4:
[0656] When the user clicks the "Start" button to start interacting with the avatar, the terminal notifies the server of this operation. Input: User operation. Output: Notification to the server.
[0657] Step 5:
[0658] The server starts a conversation generation AI to generate an initial greeting and question based on the user's basic information. The generated message is sent to the terminal and displayed to the user. Input: User's basic information. Output: Generated initial message.
[0659] Step 6:
[0660] The user responds with text or voice. This input is sent from the device to the server. Input: User response. Output: Response sent to the server.
[0661] Step 7:
[0662] The server analyzes the input content with an emotion analysis module to extract the user's emotion. The analysis result is used as input for generating an empathetic response. Input: User's response. Output: Analyzed emotion.
[0663] Step 8:
[0664] The server generates an empathetic response based on the emotion analysis results and sends the response to the device. The device displays the generated response to the user. Input: Analyzed emotion. Output: Generated empathetic response.
[0665] Step 9:
[0666] The server saves the generated conversation and emotion information in a database. The next time the conversation is held, the content of the conversation will be updated based on this information. Input: Generated conversation and emotion information. Output: Information saved in the database.
[0667] Step 10:
[0668] The server encrypts all data and stores it securely. Users can view, modify, or delete their own data at any time. Input: Information stored in the database. Output: Encrypted data.
[0669] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0670] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely, and by combining an emotion engine, recognizes the user's emotions and generates responses accordingly. The program and processing flow of this system are described in detail below.
[0671] Initial Setup and Avatar Creation
[0672] First, the user accesses the platform and enters basic information (such as name, age, hobbies and interests). The device then sends this information to the server. The server then creates a user profile based on the received basic information and calls on image-generating AI to generate an avatar. The avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[0673] Conversation initiation and emotion recognition
[0674] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, it might generate a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" In response, the user inputs text or voice, and the device sends the input to the server.
[0675] Emotion recognition by emotion engine
[0676] The server analyzes the received input content using an emotion engine to identify the user's emotion (e.g., interest, joy, sadness). The emotion engine extracts emotional signals from various inputs such as text, voice, and facial expressions to determine the user's emotional state. For example, if a user inputs, "I've been reading mystery novels recently," the emotion engine will identify "interest" and "joy."
[0677] Generating empathetic responses and interacting
[0678] Based on this analyzed emotional information, the server generates an empathetic response, such as, "That sounds interesting. Which mystery novel did you particularly enjoy?" Furthermore, the avatar's facial expression and attitude change according to this emotional information, visually demonstrating empathy.
[0679] The generated response message is displayed to the user via the terminal, and this dialogue cycle is repeated.Furthermore, the server stores the generated conversation and emotion information in a database and manages the conversation history.
[0680] Continuous personalization
[0681] The server optimizes the content of the next conversation based on the history in the database. It learns changes in the user's hobbies and interests and reflects these in the content of the conversation. For example, if the user talks about a new hobby (e.g., painting), the server continues the conversation by asking, "Regarding the painting we talked about last time, have you painted any works recently?"
[0682] Safety and Privacy Controls
[0683] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon the user's request, and hides other data. Users can view their own data within the platform and modify or delete it as necessary. The server updates the information in the database in response to this request, protecting privacy.
[0684] Specific examples
[0685] Example 1: First conversation
[0686] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0687] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0688] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0689] Example 2: Continuation Conversation
[0690] A user types, "I've been reading mystery novels lately."
[0691] The server analyzes the input using an emotion engine and determines it as "interest."
[0692] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0693] The terminal displays this message to the user.
[0694] In this way, the system provides a personalized and empathetic interaction experience for the user, reducing feelings of loneliness and leveraging the emotion engine to further address the user's emotional state.
[0695] The processing flow will be explained below.
[0696] Step 1:
[0697] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[0698] Step 2:
[0699] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[0700] Step 3:
[0701] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[0702] Step 4:
[0703] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[0704] Step 5:
[0705] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[0706] Step 6:
[0707] The device displays the generated message to the user, who can then respond with text or voice input.
[0708] Step 7:
[0709] The device sends the user's input (text or voice) to the server.
[0710] Step 8:
[0711] The server analyzes the received input using an emotion engine to identify the user's emotions, such as "joy," "interest," and "sadness."
[0712] Step 9:
[0713] The server then sends commands to the device to change the avatar's facial expression and attitude based on the extracted emotional information. For example, if the user inputs "I've been reading mystery novels recently," the avatar will display expressions and attitudes that express "interest" or "joy."
[0714] Step 10:
[0715] The server generates an empathetic response based on the emotional information, for example, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0716] Step 11:
[0717] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[0718] Step 12:
[0719] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[0720] Step 13:
[0721] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[0722] Step 14:
[0723] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[0724] Step 15:
[0725] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[0726] Specific examples
[0727] Example 1: First conversation
[0728] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0729] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0730] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0731] Example 2: Continuation Conversation
[0732] A user types, "I've been reading mystery novels lately."
[0733] The server analyzes the input using an emotion engine and determines it as "interest."
[0734] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0735] The device displays a response message, and the avatar's facial expression and attitude change to reflect interest.
[0736] Example 2
[0737] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0738] In modern society, the number of people feeling lonely is increasing, and there is a need for systems that provide empathetic dialogue tailored to individual needs. However, conventional dialogue systems have difficulty fully analyzing users' emotions and providing empathetic responses. They also fall short in terms of protecting the safety and privacy of users' personal information. A system that solves these problems and reduces feelings of loneliness is needed.
[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0740] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next dialogue based on the stored information, means for encrypting the user's personal information and safely storing the data, means for changing the avatar's facial expression and attitude based on the emotional information, means for displaying necessary information and hiding other data in response to a user request, and means for the user to view, modify, and delete data. This makes it possible to provide an empathetic dialogue that is responsive to the user's emotions, while protecting the safety and privacy of the personal information.
[0741] "User" refers to an individual who interacts with the system.
[0742] "Basic Information" refers to personal information entered by the user, such as name, age, hobbies and interests.
[0743] A "terminal" is a device that is operated by a user to input and display information.
[0744] "Server" refers to a central computer system that receives and processes information sent by users.
[0745] "Image generation artificial intelligence" refers to artificial intelligence technology for generating avatars based on basic user information.
[0746] An "avatar" is a virtual persona or character created based on a user's basic information.
[0747] "Conversational generative artificial intelligence" refers to artificial intelligence that has the technology to generate dialogue content based on user input.
[0748] An "emotion engine" refers to algorithms and technologies for analyzing and identifying emotions from user input.
[0749] "Emotion information" refers to data on the user's emotional state analyzed by the emotion engine.
[0750] "Empathetic response" refers to a dialogue that is based on analyzed emotional information and is sensitive to the user's emotions.
[0751] "Database" refers to a data structure for storing and managing generated conversation and emotion information.
[0752] "Encryption" refers to the technology of converting digital data into a form that cannot be read by third parties.
[0753] "Safe storage" refers to measures to protect users' personal information and data from unauthorized access and information leaks.
[0754] "Changing facial expressions and attitudes" refers to changing the visual response of the avatar displayed to the user based on emotional information.
[0755] "Displaying information on demand" refers to displaying only the necessary information based on the user's instructions.
[0756] "Viewing, modifying, and deleting data" refers to the actions that users can take to view, modify, or delete their own data through the platform.
[0757] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely. This system has the following configuration to recognize the user's emotions by combining emotion engines and generate responses according to those emotions.
[0758] Initial Setup and Avatar Creation
[0759] First, a user accesses the platform and enters basic information such as name, age, hobbies, and interests. The device then sends this information to the server. The server creates a user profile based on the received information and calls on image-generating AI (e.g., DeepArt) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input information.
[0760] The generated avatar is displayed to the user via the terminal and they are asked to confirm. Through this process, the user can obtain an avatar that suits their preferences.
[0761] Conversation initiation and emotion recognition
[0762] To begin a conversation with an avatar, a user clicks the "Start" button. This action causes the device to notify the server. The server then launches a conversation-generating AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[0763] For example, the server might generate a message like, "Hello, [username]. I heard you like reading. What book have you read recently?" The user can respond with text or voice, and the device will send that input to the server.
[0764] Emotion recognition by emotion engine
[0765] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer), which extracts emotional signals from inputs such as text, voice, and facial expressions to identify the user's emotions (interest, joy, sadness, etc.).
[0766] For example, if a user types, "I recently read a mystery novel," the emotion engine identifies emotions like "interest" and "joy," and based on this, the server generates a more empathetic response.
[0767] Generating empathetic responses and interacting
[0768] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds interesting. Which mystery novel did you particularly enjoy?" The avatar's facial expressions and attitude also change according to this emotional information, visually demonstrating empathy.
[0769] The generated responses are displayed to the user via the device, and this dialogue cycle is repeated. At the same time, the server stores the generated conversation and emotional information in a database and manages the conversation history. This history is used to optimize the content of the next dialogue.
[0770] Continuous personalization
[0771] The server optimizes the content of the next conversation based on the saved history data. This allows the server to learn about changes in the user's hobbies and interests and reflect that information in the conversation. For example, if the user takes up a new hobby (e.g., painting), the conversation can continue with a prompt such as, "Regarding the painting we talked about last time, have you painted any recent works?"
[0772] Safety and Privacy Controls
[0773] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon user request, hiding other data. Furthermore, users can view their own data within the platform and modify or delete it as needed. The server updates the information in the database in response to these requests, protecting privacy.
[0774] Specific examples
[0775] Example 1: First conversation
[0776] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0777] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0778] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0779] Example 2: Continuation Conversation
[0780] A user types, "I've been reading mystery novels lately."
[0781] The server analyzes the input using an emotion engine and determines it as "interest."
[0782] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[0783] The terminal displays this message to the user.
[0784] In this way, the system of the present invention can provide a personalized and empathetic interaction experience for the user, reducing feelings of loneliness. Furthermore, by utilizing the emotion engine, the system can generate appropriate responses to the user's emotional state, improving the quality of the interaction.
[0785] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0786] Step 1: Enter basic information
[0787] Users access the platform and enter basic information such as their name, age, hobbies and interests.
[0788] Input: Basic information such as name, age, hobbies and interests
[0789] Output: Basic information data
[0790] Step 2: Submit basic information
[0791] The terminal transmits the input basic information to the server.
[0792] Input: Basic information data
[0793] Output: Basic information data sent to the server
[0794] Step 3: Create your profile and create your avatar
[0795] The server creates a user profile based on the received basic information and calls an image-generating artificial intelligence (e.g., DeepArt) to generate an avatar.
[0796] Input: Basic information data sent to the server
[0797] Output: Generated avatar data
[0798] Step 4: View and check your avatar
[0799] The terminal displays the generated avatar to the user and asks for confirmation.
[0800] Input: Generated avatar data
[0801] Output: The avatar displayed to the user
[0802] Step 5: Initiating an initial conversation
[0803] The user clicks the "Get Started" button.
[0804] The terminal notifies the server of this operation.
[0805] Input: User clicks the "Get Started" button
[0806] Output: Notification to the server
[0807] Step 6: Generate initial greetings and questions
[0808] The server launches a conversational AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[0809] Input: Notification to the server, basic information data of the user
[0810] Output: Initial greeting and question message
[0811] Step 7: User response input
[0812] The user enters a response by text or voice.
[0813] The terminal transmits the input response content to the server.
[0814] Input: User response
[0815] Output: Response sent to the server
[0816] Step 8: Emotion Recognition
[0817] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the emotion.
[0818] Input: Response sent to the server
[0819] Output: Identified emotion information
[0820] Step 9: Generate an empathetic response
[0821] The server generates an empathetic response based on the analyzed emotional information, such as "That sounds interesting. Which mystery novel did you particularly enjoy?", and also changes the avatar's facial expression and attitude according to this emotional information.
[0822] Input: Identified emotion information
[0823] Output: Empathetic response message, modified avatar data
[0824] Step 10: View the response
[0825] The generated response message is displayed to the user through the terminal, along with the avatar's facial expression and attitude.
[0826] Input: Empathetic response message, modified avatar data
[0827] Output: Response message displayed to the user and the avatar's facial expression and attitude
[0828] Step 11: Saving conversation history
[0829] The server stores the generated conversation and emotion information in a database.
[0830] Input: Generated conversation data, identified emotion information
[0831] Output: Conversation data and emotion information stored in a database
[0832] Step 12: Optimize the content of your next conversation
[0833] The server optimizes the content of the next conversation based on the stored historical data.
[0834] Input: Conversation data and emotion information stored in a database
[0835] Output: Optimized next dialogue
[0836] Step 13: Managing your personal information
[0837] The server encrypts and securely stores users' personal information.
[0838] The device displays necessary information only when requested by the user, and hides other data. The user can also view, modify, and delete data.
[0839] Input: User's personal information, User's request
[0840] Output: Encrypted personal information, securely stored data, data displayed, modified, or deleted upon user request
[0841] Through these steps, it is possible to provide users with personalized, empathetic interactions and reduce their feelings of loneliness. Furthermore, by utilizing the emotion engine, it is possible to generate appropriate responses according to the user's emotional state, improving the quality of the interactions.
[0842] (Application example 2)
[0843] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0844] Conventional food delivery applications lacked a way to address the feelings of loneliness felt by users when ordering food. Furthermore, they were unable to provide empathetic dialogue that responded to the user's emotional state, making it difficult to improve the quality of the user experience.
[0845] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0846] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the next conversation content based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, means for providing empathetic dialogue if the user feels lonely based on the input content when ordering a meal, and means for analyzing the user's emotions when ordering a meal using an emotion engine and suggesting a response based on the emotions. This makes it possible to provide psychological support through empathetic dialogue even when the user feels lonely, thereby improving the quality of the user experience.
[0847] "Basic information" is personal information such as the user's name, age, hobbies and interests.
[0848] "Image generation artificial intelligence" is an artificial intelligence technology that generates images based on input data.
[0849] An "avatar" is a virtual character created to represent a user.
[0850] "Conversational artificial intelligence" is an artificial intelligence technology that generates natural dialogue based on user input.
[0851] An "emotion engine" is a technology for analyzing emotions from user input and behavior.
[0852] An "empathetic response" is a response that understands the user's emotions and is generated in a way that is sympathetic to those emotions.
[0853] The "database" is an information storage system for storing users' conversation history and emotional information.
[0854] "Encryption" is a technology that converts data to safely store users' personal information and prevents external access.
[0855] "Food delivery" is a service that allows users to order meals online and have them delivered to a specified location.
[0856] "Loneliness" is a psychological state in which one feels disconnected from others and isolated.
[0857] "Psychological support" is assistance to provide users with psychological stability and a sense of security.
[0858] This invention is a system for realizing a food delivery application that provides empathetic dialogue when a user feels lonely when ordering a meal. This system operates using a smartphone and a server.
[0859] First, a user accesses the application using their smartphone and enters basic information such as name, age, hobbies and interests, which is then sent from the device to a server, which then creates a user profile.
[0860] Next, the server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input. The generated avatar is displayed on the device and the user is asked to confirm it.
[0861] When a user starts interacting with an avatar, the server activates a conversational AI (e.g., GPT-3) to generate an initial greeting and question. Based on the user's basic information, the server generates empathetic messages such as "Hello, how busy have you been?" or "Do you have any particular favorite dishes?" The user inputs the information as text or voice, which is then sent from the device to the server.
[0862] The server analyzes the user's input using an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to identify the user's emotion. For example, if a user inputs, "I often eat alone, and I feel a little lonely," the emotion engine will identify the emotion as "loneliness."
[0863] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds lonely. What kind of person would be good to be with at a time like this?" The generated response is then displayed to the user via the device.
[0864] This cycle repeats, and the server stores the generated conversation and emotional information in a database, optimizing the content of the next conversation based on the stored information and adapting to changes in the user's hobbies and interests.
[0865] In addition, users' personal information is encrypted and securely stored by the server. This allows users to view, modify, and delete their data within the platform. The device will only display necessary information upon user request, and will hide other data.
[0866] As a concrete example, the following cases can be considered:
[0867] Example 1: First conversation
[0868] The user types, "Hello, I'm Taro Tanaka, I'm 25 years old, and my hobby is cooking."
[0869] The server generates the message "Hello, Tanaka-san. I heard you like cooking. What dish have you made recently?"
[0870] Example 2: Disconnected conversations
[0871] The user types, "I often eat alone and feel a little lonely."
[0872] The server generates an empathetic response: "That sounds lonely. Who do you think would be good to have around at a time like this?"
[0873] Prompt Sentence Examples
[0874] User: "Hi, I'm very tired."
[0875] AI: "I see, that's tough. What was particularly tiring for you today?"
[0876] In this way, the system of the present invention provides users with an empathetic interactive experience to reduce feelings of loneliness and realize psychological support.
[0877] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0878] Step 1:
[0879] The user uses a smartphone to enter basic information (name, age, hobbies, interests, etc.).
[0880] Input: Name, age, hobbies and interests
[0881] Output: Basic information set (user information completed)
[0882] What happens: A user enters the required information into a form and clicks the submit button.
[0883] Step 2:
[0884] The terminal sends the user's basic information to the server.
[0885] Input: Basic information
[0886] Output: Send data to the server
[0887] Operation: The device encrypts the basic information entered and sends it to the server.
[0888] Step 3:
[0889] The server uses image generation artificial intelligence to generate an avatar based on basic information.
[0890] Input: Encrypted basic information
[0891] Output: The generated avatar
[0892] How it works: The server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar that reflects the user's characteristics.
[0893] Step 4:
[0894] The terminal displays the generated avatar to the user and asks for confirmation.
[0895] Input: Generated avatar
[0896] Output: Display avatar and confirm user
[0897] What it does: The device displays the avatar on the screen and prompts the user to confirm the avatar.
[0898] Step 5:
[0899] The user clicks the "Start" button to begin interacting with the avatar.
[0900] Input: Start Operation
[0901] Output: Trigger to start a conversation
[0902] Action: The user clicks the "Get Started" button on the application screen.
[0903] Step 6:
[0904] The server activates a conversation-generating artificial intelligence (AI) to generate initial greetings and questions based on the user's basic information.
[0905] Input: Basic information, trigger to start a conversation
[0906] Output: Initial greetings and questions
[0907] How it works: The server uses a generative conversational AI (e.g., GPT-3) to generate messages such as "Hello, how have you been?"
[0908] Step 7:
[0909] The user inputs text or voice, and the device sends the input to the server.
[0910] Input: User text or voice input
[0911] Output: Sending user input to the server
[0912] Operation: The device sends the user's text or voice data to the server.
[0913] Step 8:
[0914] The server analyzes the user's input using an emotion engine and identifies the emotion.
[0915] Input: User text or voice input
[0916] Output: Identified emotion information
[0917] How it works: The server uses an emotion engine (e.g. IBM Watson or Microsoft Azure Emotion API) to analyze emotions from the input.
[0918] Step 9:
[0919] Based on the analyzed emotional information, the server generates an empathetic response.
[0920] Input: Emotion information
[0921] Output: Empathetic response message
[0922] How it works: The server generates an empathetic message based on the emotional information, for example, a response like "That sounds lonely."
[0923] Step 10:
[0924] The generated empathetic response is displayed to the user via the terminal.
[0925] Input: Empathetic response message
[0926] Output: The response message that is displayed to the user
[0927] Action: The terminal displays the response message on the screen.
[0928] Step 11:
[0929] The server stores the generated conversation and emotion information in a database.
[0930] Input: Conversation content, emotional information
[0931] Output: Save to database
[0932] How it works: The server records conversation history and emotion information in a database.
[0933] Step 12:
[0934] The conversation will be updated based on the saved information the next time you interact, and will also respond to changes in the user's hobbies and interests.
[0935] Input: saved conversation history, emotional information
[0936] Output: Optimized dialogue
[0937] How it works: The server retrieves information from the database and personalizes your next interaction.
[0938] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0939] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0940] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0941] [Third embodiment]
[0942] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0943] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0944] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0945] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0946] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0947] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0948] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0949] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0950] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0951] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0952] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0953] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0954] The present invention is a system that provides personalized, empathetic dialogue to users who feel lonely. This system uses artificial intelligence to prompt users to enter basic information and engage in dialogue through a generated avatar. It also analyzes the user's emotions and generates responses based on those emotions. The program for this system is described in detail below.
[0955] Initial Setup and Avatar Creation
[0956] First, a user accesses the platform and enters basic information (name, age, hobbies and interests). This basic information is sent from the device to the server, which then creates a user profile based on it. Next, the server uses image generation artificial intelligence to generate an avatar for the user. This avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[0957] Conversation initiation and emotional understanding
[0958] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" can be generated. In response, the user inputs text or voice, and the device sends the input to the server.
[0959] The server analyzes this input using an emotion analysis module to extract the user's emotions (e.g., interest, joy). It then generates an empathetic response based on the analyzed emotions. For example, if a user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via their device.
[0960] Continuous personalization
[0961] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This allows the server to have basic data to better personalize the content of the next conversation. For example, if the user talks about a new hobby (e.g., painting), this information will be reflected in future conversations. The conversation can continue in the form of, "Regarding the painting we talked about last time, have you painted any works recently?"
[0962] Safety and Privacy Controls
[0963] Users' personal information is encrypted and stored securely by the server. The device displays necessary information only upon the user's request, and hides other data. Users have the right to view, modify, or delete their own data within the platform, and the server updates the information in the database accordingly.
[0964] Specific examples
[0965] Example 1: First conversation
[0966] The user enters "Yamada Taro, 30 years old, hobby is reading."
[0967] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[0968] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[0969] Example 2: Continuation Conversation
[0970] A user types, "I've been reading mystery novels lately."
[0971] The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[0972] The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?"
[0973] The terminal displays this message to the user.
[0974] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[0975] The processing flow will be explained below.
[0976] Step 1:
[0977] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[0978] Step 2:
[0979] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[0980] Step 3:
[0981] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[0982] Step 4:
[0983] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[0984] Step 5:
[0985] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[0986] Step 6:
[0987] The device displays the generated message to the user, who can then respond with text or voice input.
[0988] Step 7:
[0989] The device sends the user's input (text or voice) to the server.
[0990] Step 8:
[0991] The server analyzes the received input using an emotion analysis module to identify the user's emotions, such as "joy," "interest," and "sadness."
[0992] Step 9:
[0993] The server generates an empathetic response based on the extracted emotions. For example, if a user types, "I've been reading mystery novels lately," it generates a response like, "That sounds interesting. Which mystery novels did you enjoy the most?"
[0994] Step 10:
[0995] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[0996] Step 11:
[0997] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[0998] Step 12:
[0999] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[1000] Step 13:
[1001] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[1002] Step 14:
[1003] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[1004] Example 1
[1005] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1006] In today's world, there is a need to provide users who feel lonely with an effective and empathetic dialogue experience. Conventional systems have struggled to properly recognize a user's emotions and generate personalized responses. Furthermore, they lacked a means to continuously personalize the dialogue content while safely managing user information. The present invention aims to solve these problems and provide users with human-like empathy.
[1007] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1008] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, terminal means for a user to access the platform and input basic information, means for transmitting the input information to the server and having the server create a profile, means for the terminal to display the generated avatar to the user and request confirmation, means for the terminal to notify the server of a user's dialogue start operation, means for transmitting user responses during the conversation from the terminal to the server, and means for generating an appropriate empathetic response based on the emotion analysis results. This enables the generation of an appropriate response according to the user's emotions and provides a personalized empathetic dialogue experience.
[1009] "Basic information" refers to information such as name, age, hobbies and interests that users enter into the system.
[1010] "Image generation artificial intelligence" is an artificial intelligence technology for generating avatars based on basic information about users.
[1011] An "avatar" is a digital representation with a character that is generated based on basic information about the user.
[1012] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates dialogue content based on user input.
[1013] "Emotion analysis" is a technology that analyzes emotions from user input and extracts meaning.
[1014] An "empathetic response" is a response that is generated in a way that is in tune with the user's emotions, based on analyzed emotional information.
[1015] The "database" is an information system that stores the generated conversation and emotion information and uses it to update the content of future conversations.
[1016] "Encryption" is a data protection method used to safely store users' personal information.
[1017] The present invention provides a system for providing personalized empathetic interaction to users who feel lonely. The system is implemented according to the following steps.
[1018] Initial Setup and Avatar Creation
[1019] Users access the platform via a web browser or a dedicated app. After accessing the platform, users enter basic information (name, age, hobbies and interests) into the displayed input form. This information is then sent from the device to the server.
[1020] The server creates a user profile based on the received basic information. Using this profile information, it activates an image generation AI (e.g., Stable Diffusion) to generate an avatar corresponding to the user. The generated avatar is sent to the terminal and displayed to the user. The user receives a display to confirm this avatar.
[1021] Conversation initiation and emotional understanding
[1022] When the user clicks the "Start" button, the operation is notified from the terminal to the server.
[1023] The server launches a conversational AI (e.g., GPT-4) to generate an initial dialogue message based on the user's basic information. For example, a message like "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" is generated. The device displays the message to the user. When the user responds with text or voice, that information is sent from the device to the server.
[1024] The server analyzes the input information using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). Through this analysis, the server identifies the user's emotions (e.g., interest, joy). Based on the analysis results, the server generates an appropriate empathetic response. For example, if the user responds, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. What are your favorite mystery novels?" This response is displayed to the user via their device.
[1025] Continuous personalization
[1026] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This generates basic data to further personalize the content of the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information is stored in the database and reflected in future conversations. In the next conversation, a personalized message such as, "Regarding the painting we talked about last time, have you painted any recent works?" is generated.
[1027] Safety and Privacy Controls
[1028] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users have the right to view, correct, or delete their own data within the platform, and the server updates the information in the database accordingly.
[1029] Examples and prompts
[1030] Example 1: First conversation
[1031] 1. The user enters "Yamada Taro, 30 years old, hobby is reading."
[1032] 2. The server creates a profile for "Yamada Taro" and activates an image-generating AI (e.g., Stable Diffusion) to generate an avatar.
[1033] 3. The device displays the avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1034] Example prompt sentence:
[1035] Prompt for generative AI models (image-generating AI):
[1036] "Please create an avatar for a 30-year-old Japanese man whose hobby is reading, with a kind and intelligent face."
[1037] Prompts for generative AI models (generative conversational AI):
[1038] "The user's name is Taro Yamada, and his hobby is reading. As an initial greeting to him, generate a message that says, 'Hello, Taro Yamada. I heard you like reading. What book have you read recently?'"
[1039] Example 2: Continuation Conversation
[1040] 1. The user types, "I've been reading mystery novels lately."
[1041] 2. The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[1042] 3. The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?" and displays it on the device.
[1043] Example prompt sentence:
[1044] Prompt to the generative AI model (sentiment analysis module):
[1045] "Text: I recently read a mystery novel. Please analyze my emotions and tell me what would be an appropriate response."
[1046] Prompts for generative AI models (generative conversational AI):
[1047] "A user types, 'I've been reading mystery novels lately.' Generate an empathetic response like, 'That sounds interesting. Which mystery novels did you enjoy the most?'"
[1048] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[1049] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1050] Step 1:
[1051] A user accesses the platform and enters basic information (name, age, hobbies and interests). The basic information entered by the user is sent from the device to the server. The input data includes the user's name (e.g., Yamada Taro), age (e.g., 30 years old), hobbies (e.g., reading), etc. This data is used to create the user's profile.
[1052] Input: User's basic information (name, age, hobbies and interests)
[1053] Output: Basic information sent to the server
[1054] Step 2:
[1055] The server creates a user profile based on the received basic information. This includes storing the basic information in a database. It then invokes an image-generating AI (e.g., Stable Diffusion) to generate an avatar for the user. The generated avatar has characteristics based on the basic information, such as an intelligent and kind appearance. The generated avatar is sent from the server to the device and displayed to the user. It is displayed along with the message, "Are you sure you want this avatar?"
[1056] Input: Basic information sent to the server
[1057] Output: An avatar image and a confirmation message that will be displayed to the user.
[1058] Step 3:
[1059] When the user clicks the "Start" button, the device notifies the server of the operation. The server receives this notification and launches a conversation-generating AI (e.g., GPT-4). The server generates an initial greeting and question based on the user's basic information. For example, it generates a message like, "Hello, Yamada-san. I heard you like reading. What book have you read recently?" This message is sent from the server to the device and displayed to the user.
[1060] Input: User interaction start operation
[1061] Output: Initial greeting message and question
[1062] Step 4:
[1063] When the user responds with text or voice, the input is sent from the device to the server. The server analyzes the input using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). The analysis identifies the user's emotion (e.g., interest, joy). The extracted emotion information is used in the next step.
[1064] Input: User response (text or voice)
[1065] Output: Parsed emotion data
[1066] Step 5:
[1067] The server generates an appropriate empathetic response based on the results of the emotion analysis. It uses conversation-generating AI to create an empathetic message. For example, a response might be generated such as, "That sounds interesting. What are your favorite mystery novels?" This response is sent from the server to the device and displayed to the user.
[1068] Input: Parsed emotion data
[1069] Output: Empathetic response message
[1070] Step 6:
[1071] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This stored data is used to personalize the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information will be reflected in the next conversation. In the next conversation, a personalized message will be generated, such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[1072] Input: Conversation history and emotional information
[1073] Output: Database information used for future personalization
[1074] Step 7:
[1075] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users can view, modify, and delete their own data within the platform, and the server updates the information in the database accordingly.
[1076] Input: User's personal information
[1077] Output: Encrypted database information and the ability to view, modify, and delete data according to user requests
[1078] (Application example 1)
[1079] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1080] In modern society, the number of users who feel lonely is increasing. To improve customer experience, particularly in brick-and-mortar stores, personalized, empathetic dialogue is needed, rather than simply providing guidance. However, current in-store guidance systems and guide robots have difficulty responding flexibly based on the customer's basic information and emotions. Therefore, the present invention aims to provide a system that reduces customers' feelings of loneliness and improves customer experience by providing personalized, empathetic dialogue.
[1081] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1082] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for providing personalized guidance through continuous dialogue based on the user's basic information, means for implementing the system in a store using a robot, and means for encrypting the user's personal information and securely storing the avatar and its data. This makes it possible to provide personalized empathetic dialogue in a store based on the user's basic information, hobbies, and interests.
[1083] "Basic information" refers to the initial data needed to individualize and personalize the interaction, such as the user's name, age, hobbies, and interests.
[1084] "Image generation artificial intelligence" is an artificial intelligence technology for generating an avatar corresponding to a user based on basic information about the user.
[1085] An "avatar" is a digital person or character with a virtual appearance and personality that is generated based on basic information about the user.
[1086] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates appropriate dialogue content based on input information from the user.
[1087] "Emotion analysis" is a technology for extracting emotions from user input and determining emotions such as interest, joy, sadness, etc.
[1088] An "empathetic response" is a response that is generated based on the results of analyzing the user's emotions and is given in a manner that empathizes with the user's emotions.
[1089] The "database" is an information management system that stores the generated conversation and emotion information and updates or references it as needed.
[1090] "Encryption" is the technology of converting data into cryptographic code to protect users' personal information.
[1091] "Individualized guidance" refers to guidance information tailored to the individual needs and interests of a user, provided based on the user's basic information and past interaction history.
[1092] A "robot" is a mechanical device that exists physically and has a guidance function to operate the above system in the real world.
[1093] To implement the present invention, the following system configuration and processing procedures are basically required.
[1094] A user first accesses the system and enters their basic information, including name, age, hobbies, and interests. The system then creates a user profile and uses image generation AI to generate an avatar for the user. The terminal then displays the generated avatar to the user and asks for their confirmation.
[1095] When the user clicks the "Start" button to begin a conversation with the avatar, the device notifies the server of this action. The server then activates a conversation-generating AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" may be generated.
[1096] When the user responds with text or voice, the device sends the input to the server. The server analyzes the input using an emotion analysis module to extract the user's emotion. An empathetic response is generated based on the analyzed emotion. For example, if the user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via the device.
[1097] This system realizes continuous personalized dialogue by storing the content of the dialogue with the user and emotional information in a database. The next time the dialogue is held, the content is updated based on the stored data, providing a more personalized response. For example, if the user mentioned a new hobby (e.g., painting) in the previous dialogue, the next dialogue could continue with a question such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[1098] The system encrypts users' personal information and stores avatars and related data securely, protecting users' privacy. Users can also view, modify, and delete their own data within the platform.
[1099] As a concrete example, the following system usage scenarios can be considered:
[1100] Example 1: When a guide robot approaches a customer in a store
[1101] "Hello, Mr. / Ms. XX. Thank you for coming. I heard you enjoy reading. What book have you read recently?"
[1102] An example of a prompt is:
[1103] User: Yamada Taro. Age: 30. Hobbies: reading and watching movies.
[1104] What mystery novel have you read recently?”
[1105] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[1106] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1107] Step 1:
[1108] A user accesses the system and enters their basic information. This basic information includes name, age, hobbies, interests, etc. The entered basic information is sent from the terminal to the server. Input: User's basic information. Output: Basic information saved on the server.
[1109] Step 2:
[1110] The server creates a user profile based on the received basic information. An image generation AI uses this user profile to generate an avatar. The generated avatar is sent to the device. Input: Basic user information. Output: Generated avatar.
[1111] Step 3:
[1112] The device displays the generated avatar to the user and asks for confirmation. Input: Generated avatar. Output: Displayed avatar and user confirmation.
[1113] Step 4:
[1114] When the user clicks the "Start" button to start interacting with the avatar, the terminal notifies the server of this operation. Input: User operation. Output: Notification to the server.
[1115] Step 5:
[1116] The server starts a conversation generation AI to generate an initial greeting and question based on the user's basic information. The generated message is sent to the terminal and displayed to the user. Input: User's basic information. Output: Generated initial message.
[1117] Step 6:
[1118] The user responds with text or voice. This input is sent from the device to the server. Input: User response. Output: Response sent to the server.
[1119] Step 7:
[1120] The server analyzes the input content with an emotion analysis module to extract the user's emotion. The analysis result is used as input for generating an empathetic response. Input: User's response. Output: Analyzed emotion.
[1121] Step 8:
[1122] The server generates an empathetic response based on the emotion analysis results and sends the response to the device. The device displays the generated response to the user. Input: Analyzed emotion. Output: Generated empathetic response.
[1123] Step 9:
[1124] The server saves the generated conversation and emotion information in a database. The next time the conversation is held, the content of the conversation will be updated based on this information. Input: Generated conversation and emotion information. Output: Information saved in the database.
[1125] Step 10:
[1126] The server encrypts all data and stores it securely. Users can view, modify, or delete their own data at any time. Input: Information stored in the database. Output: Encrypted data.
[1127] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1128] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely, and by combining an emotion engine, recognizes the user's emotions and generates responses accordingly. The program and processing flow of this system are described in detail below.
[1129] Initial Setup and Avatar Creation
[1130] First, the user accesses the platform and enters basic information (such as name, age, hobbies and interests). The device then sends this information to the server. The server then creates a user profile based on the received basic information and calls on image-generating AI to generate an avatar. The avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[1131] Conversation initiation and emotion recognition
[1132] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, it might generate a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" In response, the user inputs text or voice, and the device sends the input to the server.
[1133] Emotion recognition by emotion engine
[1134] The server analyzes the received input content using an emotion engine to identify the user's emotion (e.g., interest, joy, sadness). The emotion engine extracts emotional signals from various inputs such as text, voice, and facial expressions to determine the user's emotional state. For example, if a user inputs, "I've been reading mystery novels recently," the emotion engine will identify "interest" and "joy."
[1135] Generating empathetic responses and interacting
[1136] Based on this analyzed emotional information, the server generates an empathetic response, such as, "That sounds interesting. Which mystery novel did you particularly enjoy?" Furthermore, the avatar's facial expression and attitude change according to this emotional information, visually demonstrating empathy.
[1137] The generated response message is displayed to the user via the terminal, and this dialogue cycle is repeated.Furthermore, the server stores the generated conversation and emotion information in a database and manages the conversation history.
[1138] Continuous personalization
[1139] The server optimizes the content of the next conversation based on the history in the database. It learns changes in the user's hobbies and interests and reflects these in the content of the conversation. For example, if the user talks about a new hobby (e.g., painting), the server continues the conversation by asking, "Regarding the painting we talked about last time, have you painted any works recently?"
[1140] Safety and Privacy Controls
[1141] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon the user's request, and hides other data. Users can view their own data within the platform and modify or delete it as necessary. The server updates the information in the database in response to this request, protecting privacy.
[1142] Specific examples
[1143] Example 1: First conversation
[1144] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1145] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1146] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1147] Example 2: Continuation Conversation
[1148] A user types, "I've been reading mystery novels lately."
[1149] The server analyzes the input using an emotion engine and determines it as "interest."
[1150] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1151] The terminal displays this message to the user.
[1152] In this way, the system provides a personalized and empathetic interaction experience for the user, reducing feelings of loneliness and leveraging the emotion engine to further address the user's emotional state.
[1153] The processing flow will be explained below.
[1154] Step 1:
[1155] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[1156] Step 2:
[1157] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[1158] Step 3:
[1159] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[1160] Step 4:
[1161] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[1162] Step 5:
[1163] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[1164] Step 6:
[1165] The device displays the generated message to the user, who can then respond with text or voice input.
[1166] Step 7:
[1167] The device sends the user's input (text or voice) to the server.
[1168] Step 8:
[1169] The server analyzes the received input using an emotion engine to identify the user's emotions, such as "joy," "interest," and "sadness."
[1170] Step 9:
[1171] The server then sends commands to the device to change the avatar's facial expression and attitude based on the extracted emotional information. For example, if the user inputs "I've been reading mystery novels recently," the avatar will display expressions and attitudes that express "interest" or "joy."
[1172] Step 10:
[1173] The server generates an empathetic response based on the emotional information, for example, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1174] Step 11:
[1175] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[1176] Step 12:
[1177] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[1178] Step 13:
[1179] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[1180] Step 14:
[1181] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[1182] Step 15:
[1183] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[1184] Specific examples
[1185] Example 1: First conversation
[1186] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1187] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1188] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1189] Example 2: Continuation Conversation
[1190] A user types, "I've been reading mystery novels lately."
[1191] The server analyzes the input using an emotion engine and determines it as "interest."
[1192] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1193] The device displays a response message, and the avatar's facial expression and attitude change to reflect interest.
[1194] Example 2
[1195] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1196] In modern society, the number of people feeling lonely is increasing, and there is a need for systems that provide empathetic dialogue tailored to individual needs. However, conventional dialogue systems have difficulty fully analyzing users' emotions and providing empathetic responses. They also fall short in terms of protecting the safety and privacy of users' personal information. A system that solves these problems and reduces feelings of loneliness is needed.
[1197] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1198] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next dialogue based on the stored information, means for encrypting the user's personal information and safely storing the data, means for changing the avatar's facial expression and attitude based on the emotional information, means for displaying necessary information and hiding other data in response to a user request, and means for the user to view, modify, and delete data. This makes it possible to provide an empathetic dialogue that is responsive to the user's emotions, while protecting the safety and privacy of the personal information.
[1199] "User" refers to an individual who interacts with the system.
[1200] "Basic Information" refers to personal information entered by the user, such as name, age, hobbies and interests.
[1201] A "terminal" is a device that is operated by a user to input and display information.
[1202] "Server" refers to a central computer system that receives and processes information sent by users.
[1203] "Image generation artificial intelligence" refers to artificial intelligence technology for generating avatars based on basic user information.
[1204] An "avatar" is a virtual persona or character created based on a user's basic information.
[1205] "Conversational generative artificial intelligence" refers to artificial intelligence that has the technology to generate dialogue content based on user input.
[1206] An "emotion engine" refers to algorithms and technologies for analyzing and identifying emotions from user input.
[1207] "Emotion information" refers to data on the user's emotional state analyzed by the emotion engine.
[1208] "Empathetic response" refers to a dialogue that is based on analyzed emotional information and is sensitive to the user's emotions.
[1209] "Database" refers to a data structure for storing and managing generated conversation and emotion information.
[1210] "Encryption" refers to the technology of converting digital data into a form that cannot be read by third parties.
[1211] "Safe storage" refers to measures to protect users' personal information and data from unauthorized access and information leaks.
[1212] "Changing facial expressions and attitudes" refers to changing the visual response of the avatar displayed to the user based on emotional information.
[1213] "Displaying information on demand" refers to displaying only the necessary information based on the user's instructions.
[1214] "Viewing, modifying, and deleting data" refers to the actions that users can take to view, modify, or delete their own data through the platform.
[1215] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely. This system has the following configuration to recognize the user's emotions by combining emotion engines and generate responses according to those emotions.
[1216] Initial Setup and Avatar Creation
[1217] First, a user accesses the platform and enters basic information such as name, age, hobbies, and interests. The device then sends this information to the server. The server creates a user profile based on the received information and calls on image-generating AI (e.g., DeepArt) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input information.
[1218] The generated avatar is displayed to the user via the terminal and they are asked to confirm. Through this process, the user can obtain an avatar that suits their preferences.
[1219] Conversation initiation and emotion recognition
[1220] To begin a conversation with an avatar, a user clicks the "Start" button. This action causes the device to notify the server. The server then launches a conversation-generating AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[1221] For example, the server might generate a message like, "Hello, [username]. I heard you like reading. What book have you read recently?" The user can respond with text or voice, and the device will send that input to the server.
[1222] Emotion recognition by emotion engine
[1223] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer), which extracts emotional signals from inputs such as text, voice, and facial expressions to identify the user's emotions (interest, joy, sadness, etc.).
[1224] For example, if a user types, "I recently read a mystery novel," the emotion engine identifies emotions like "interest" and "joy," and based on this, the server generates a more empathetic response.
[1225] Generating empathetic responses and interacting
[1226] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds interesting. Which mystery novel did you particularly enjoy?" The avatar's facial expressions and attitude also change according to this emotional information, visually demonstrating empathy.
[1227] The generated responses are displayed to the user via the device, and this dialogue cycle is repeated. At the same time, the server stores the generated conversation and emotional information in a database and manages the conversation history. This history is used to optimize the content of the next dialogue.
[1228] Continuous personalization
[1229] The server optimizes the content of the next conversation based on the saved history data. This allows the server to learn about changes in the user's hobbies and interests and reflect that information in the conversation. For example, if the user takes up a new hobby (e.g., painting), the conversation can continue with a prompt such as, "Regarding the painting we talked about last time, have you painted any recent works?"
[1230] Safety and Privacy Controls
[1231] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon user request, hiding other data. Furthermore, users can view their own data within the platform and modify or delete it as needed. The server updates the information in the database in response to these requests, protecting privacy.
[1232] Specific examples
[1233] Example 1: First conversation
[1234] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1235] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1236] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1237] Example 2: Continuation Conversation
[1238] A user types, "I've been reading mystery novels lately."
[1239] The server analyzes the input using an emotion engine and determines it as "interest."
[1240] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1241] The terminal displays this message to the user.
[1242] In this way, the system of the present invention can provide a personalized and empathetic interaction experience for the user, reducing feelings of loneliness. Furthermore, by utilizing the emotion engine, the system can generate appropriate responses to the user's emotional state, improving the quality of the interaction.
[1243] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1244] Step 1: Enter basic information
[1245] Users access the platform and enter basic information such as their name, age, hobbies and interests.
[1246] Input: Basic information such as name, age, hobbies and interests
[1247] Output: Basic information data
[1248] Step 2: Submit basic information
[1249] The terminal transmits the input basic information to the server.
[1250] Input: Basic information data
[1251] Output: Basic information data sent to the server
[1252] Step 3: Create your profile and create your avatar
[1253] The server creates a user profile based on the received basic information and calls an image-generating artificial intelligence (e.g., DeepArt) to generate an avatar.
[1254] Input: Basic information data sent to the server
[1255] Output: Generated avatar data
[1256] Step 4: View and check your avatar
[1257] The terminal displays the generated avatar to the user and asks for confirmation.
[1258] Input: Generated avatar data
[1259] Output: The avatar displayed to the user
[1260] Step 5: Initiating an initial conversation
[1261] The user clicks the "Get Started" button.
[1262] The terminal notifies the server of this operation.
[1263] Input: User clicks the "Get Started" button
[1264] Output: Notification to the server
[1265] Step 6: Generate initial greetings and questions
[1266] The server launches a conversational AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[1267] Input: Notification to the server, basic information data of the user
[1268] Output: Initial greeting and question message
[1269] Step 7: User response input
[1270] The user enters a response by text or voice.
[1271] The terminal transmits the input response content to the server.
[1272] Input: User response
[1273] Output: Response sent to the server
[1274] Step 8: Emotion Recognition
[1275] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the emotion.
[1276] Input: Response sent to the server
[1277] Output: Identified emotion information
[1278] Step 9: Generate an empathetic response
[1279] The server generates an empathetic response based on the analyzed emotional information, such as "That sounds interesting. Which mystery novel did you particularly enjoy?", and also changes the avatar's facial expression and attitude according to this emotional information.
[1280] Input: Identified emotion information
[1281] Output: Empathetic response message, modified avatar data
[1282] Step 10: View the response
[1283] The generated response message is displayed to the user through the terminal, along with the avatar's facial expression and attitude.
[1284] Input: Empathetic response message, modified avatar data
[1285] Output: Response message displayed to the user and the avatar's facial expression and attitude
[1286] Step 11: Saving conversation history
[1287] The server stores the generated conversation and emotion information in a database.
[1288] Input: Generated conversation data, identified emotion information
[1289] Output: Conversation data and emotion information stored in a database
[1290] Step 12: Optimize the content of your next conversation
[1291] The server optimizes the content of the next conversation based on the stored historical data.
[1292] Input: Conversation data and emotion information stored in a database
[1293] Output: Optimized next dialogue
[1294] Step 13: Managing your personal information
[1295] The server encrypts and securely stores users' personal information.
[1296] The device displays necessary information only when requested by the user, and hides other data. The user can also view, modify, and delete data.
[1297] Input: User's personal information, User's request
[1298] Output: Encrypted personal information, securely stored data, data displayed, modified, or deleted upon user request
[1299] Through these steps, it is possible to provide users with personalized, empathetic interactions and reduce their feelings of loneliness. Furthermore, by utilizing the emotion engine, it is possible to generate appropriate responses according to the user's emotional state, improving the quality of the interactions.
[1300] (Application example 2)
[1301] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1302] Conventional food delivery applications lacked a way to address the feelings of loneliness felt by users when ordering food. Furthermore, they were unable to provide empathetic dialogue that responded to the user's emotional state, making it difficult to improve the quality of the user experience.
[1303] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1304] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the next conversation content based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, means for providing empathetic dialogue if the user feels lonely based on the input content when ordering a meal, and means for analyzing the user's emotions when ordering a meal using an emotion engine and suggesting a response based on the emotions. This makes it possible to provide psychological support through empathetic dialogue even when the user feels lonely, thereby improving the quality of the user experience.
[1305] "Basic information" is personal information such as the user's name, age, hobbies and interests.
[1306] "Image generation artificial intelligence" is an artificial intelligence technology that generates images based on input data.
[1307] An "avatar" is a virtual character created to represent a user.
[1308] "Conversational artificial intelligence" is an artificial intelligence technology that generates natural dialogue based on user input.
[1309] An "emotion engine" is a technology for analyzing emotions from user input and behavior.
[1310] An "empathetic response" is a response that understands the user's emotions and is generated in a way that is sympathetic to those emotions.
[1311] The "database" is an information storage system for storing users' conversation history and emotional information.
[1312] "Encryption" is a technology that converts data to safely store users' personal information and prevents external access.
[1313] "Food delivery" is a service that allows users to order meals online and have them delivered to a specified location.
[1314] "Loneliness" is a psychological state in which one feels disconnected from others and isolated.
[1315] "Psychological support" is assistance to provide users with psychological stability and a sense of security.
[1316] This invention is a system for realizing a food delivery application that provides empathetic dialogue when a user feels lonely when ordering a meal. This system operates using a smartphone and a server.
[1317] First, a user accesses the application using their smartphone and enters basic information such as name, age, hobbies and interests, which is then sent from the device to a server, which then creates a user profile.
[1318] Next, the server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input. The generated avatar is displayed on the device and the user is asked to confirm it.
[1319] When a user starts interacting with an avatar, the server activates a conversational AI (e.g., GPT-3) to generate an initial greeting and question. Based on the user's basic information, the server generates empathetic messages such as "Hello, how busy have you been?" or "Do you have any particular favorite dishes?" The user inputs the information as text or voice, which is then sent from the device to the server.
[1320] The server analyzes the user's input using an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to identify the user's emotion. For example, if a user inputs, "I often eat alone, and I feel a little lonely," the emotion engine will identify the emotion as "loneliness."
[1321] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds lonely. What kind of person would be good to be with at a time like this?" The generated response is then displayed to the user via the device.
[1322] This cycle repeats, and the server stores the generated conversation and emotional information in a database, optimizing the content of the next conversation based on the stored information and adapting to changes in the user's hobbies and interests.
[1323] In addition, users' personal information is encrypted and securely stored by the server. This allows users to view, modify, and delete their data within the platform. The device will only display necessary information upon user request, and will hide other data.
[1324] As a concrete example, the following cases can be considered:
[1325] Example 1: First conversation
[1326] The user types, "Hello, I'm Taro Tanaka, I'm 25 years old, and my hobby is cooking."
[1327] The server generates the message "Hello, Tanaka-san. I heard you like cooking. What dish have you made recently?"
[1328] Example 2: Disconnected conversations
[1329] The user types, "I often eat alone and feel a little lonely."
[1330] The server generates an empathetic response: "That sounds lonely. Who do you think would be good to have around at a time like this?"
[1331] Prompt Sentence Examples
[1332] User: "Hi, I'm very tired."
[1333] AI: "I see, that's tough. What was particularly tiring for you today?"
[1334] In this way, the system of the present invention provides users with an empathetic interactive experience to reduce feelings of loneliness and realize psychological support.
[1335] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1336] Step 1:
[1337] The user uses a smartphone to enter basic information (name, age, hobbies, interests, etc.).
[1338] Input: Name, age, hobbies and interests
[1339] Output: Basic information set (user information completed)
[1340] What happens: A user enters the required information into a form and clicks the submit button.
[1341] Step 2:
[1342] The terminal sends the user's basic information to the server.
[1343] Input: Basic information
[1344] Output: Send data to the server
[1345] Operation: The device encrypts the basic information entered and sends it to the server.
[1346] Step 3:
[1347] The server uses image generation artificial intelligence to generate an avatar based on basic information.
[1348] Input: Encrypted basic information
[1349] Output: The generated avatar
[1350] How it works: The server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar that reflects the user's characteristics.
[1351] Step 4:
[1352] The terminal displays the generated avatar to the user and asks for confirmation.
[1353] Input: Generated avatar
[1354] Output: Display avatar and confirm user
[1355] What it does: The device displays the avatar on the screen and prompts the user to confirm the avatar.
[1356] Step 5:
[1357] The user clicks the "Start" button to begin interacting with the avatar.
[1358] Input: Start Operation
[1359] Output: Trigger to start a conversation
[1360] Action: The user clicks the "Get Started" button on the application screen.
[1361] Step 6:
[1362] The server activates a conversation-generating artificial intelligence (AI) to generate initial greetings and questions based on the user's basic information.
[1363] Input: Basic information, trigger to start a conversation
[1364] Output: Initial greetings and questions
[1365] How it works: The server uses a generative conversational AI (e.g., GPT-3) to generate messages such as "Hello, how have you been?"
[1366] Step 7:
[1367] The user inputs text or voice, and the device sends the input to the server.
[1368] Input: User text or voice input
[1369] Output: Sending user input to the server
[1370] Operation: The device sends the user's text or voice data to the server.
[1371] Step 8:
[1372] The server analyzes the user's input using an emotion engine and identifies the emotion.
[1373] Input: User text or voice input
[1374] Output: Identified emotion information
[1375] How it works: The server uses an emotion engine (e.g. IBM Watson or Microsoft Azure Emotion API) to analyze emotions from the input.
[1376] Step 9:
[1377] Based on the analyzed emotional information, the server generates an empathetic response.
[1378] Input: Emotion information
[1379] Output: Empathetic response message
[1380] How it works: The server generates an empathetic message based on the emotional information, for example, a response like "That sounds lonely."
[1381] Step 10:
[1382] The generated empathetic response is displayed to the user via the terminal.
[1383] Input: Empathetic response message
[1384] Output: The response message that is displayed to the user
[1385] Action: The terminal displays the response message on the screen.
[1386] Step 11:
[1387] The server stores the generated conversation and emotion information in a database.
[1388] Input: Conversation content, emotional information
[1389] Output: Save to database
[1390] How it works: The server records conversation history and emotion information in a database.
[1391] Step 12:
[1392] The conversation will be updated based on the saved information the next time you interact, and will also respond to changes in the user's hobbies and interests.
[1393] Input: saved conversation history, emotional information
[1394] Output: Optimized dialogue
[1395] How it works: The server retrieves information from the database and personalizes your next interaction.
[1396] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1398] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1399] [Fourth embodiment]
[1400] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1401] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1402] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1403] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1404] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1406] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1407] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1408] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1409] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1410] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1411] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1412] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1413] The present invention is a system that provides personalized, empathetic dialogue to users who feel lonely. This system uses artificial intelligence to prompt users to enter basic information and engage in dialogue through a generated avatar. It also analyzes the user's emotions and generates responses based on those emotions. The program for this system is described in detail below.
[1414] Initial Setup and Avatar Creation
[1415] First, a user accesses the platform and enters basic information (name, age, hobbies and interests). This basic information is sent from the device to the server, which then creates a user profile based on it. Next, the server uses image generation artificial intelligence to generate an avatar for the user. This avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[1416] Conversation initiation and emotional understanding
[1417] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" can be generated. In response, the user inputs text or voice, and the device sends the input to the server.
[1418] The server analyzes this input using an emotion analysis module to extract the user's emotions (e.g., interest, joy). It then generates an empathetic response based on the analyzed emotions. For example, if a user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via their device.
[1419] Continuous personalization
[1420] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This allows the server to have basic data to better personalize the content of the next conversation. For example, if the user talks about a new hobby (e.g., painting), this information will be reflected in future conversations. The conversation can continue in the form of, "Regarding the painting we talked about last time, have you painted any works recently?"
[1421] Safety and Privacy Controls
[1422] Users' personal information is encrypted and stored securely by the server. The device displays necessary information only upon the user's request, and hides other data. Users have the right to view, modify, or delete their own data within the platform, and the server updates the information in the database accordingly.
[1423] Specific examples
[1424] Example 1: First conversation
[1425] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1426] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1427] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1428] Example 2: Continuation Conversation
[1429] A user types, "I've been reading mystery novels lately."
[1430] The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[1431] The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?"
[1432] The terminal displays this message to the user.
[1433] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[1434] The processing flow will be explained below.
[1435] Step 1:
[1436] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[1437] Step 2:
[1438] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[1439] Step 3:
[1440] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[1441] Step 4:
[1442] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[1443] Step 5:
[1444] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[1445] Step 6:
[1446] The device displays the generated message to the user, who can then respond with text or voice input.
[1447] Step 7:
[1448] The device sends the user's input (text or voice) to the server.
[1449] Step 8:
[1450] The server analyzes the received input using an emotion analysis module to identify the user's emotions, such as "joy," "interest," and "sadness."
[1451] Step 9:
[1452] The server generates an empathetic response based on the extracted emotions. For example, if a user types, "I've been reading mystery novels lately," it generates a response like, "That sounds interesting. Which mystery novels did you enjoy the most?"
[1453] Step 10:
[1454] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[1455] Step 11:
[1456] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[1457] Step 12:
[1458] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[1459] Step 13:
[1460] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[1461] Step 14:
[1462] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[1463] Example 1
[1464] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1465] In today's world, there is a need to provide users who feel lonely with an effective and empathetic dialogue experience. Conventional systems have struggled to properly recognize a user's emotions and generate personalized responses. Furthermore, they lacked a means to continuously personalize the dialogue content while safely managing user information. The present invention aims to solve these problems and provide users with human-like empathy.
[1466] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1467] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, terminal means for a user to access the platform and input basic information, means for transmitting the input information to the server and having the server create a profile, means for the terminal to display the generated avatar to the user and request confirmation, means for the terminal to notify the server of a user's dialogue start operation, means for transmitting user responses during the conversation from the terminal to the server, and means for generating an appropriate empathetic response based on the emotion analysis results. This enables the generation of an appropriate response according to the user's emotions and provides a personalized empathetic dialogue experience.
[1468] "Basic information" refers to information such as name, age, hobbies and interests that users enter into the system.
[1469] "Image generation artificial intelligence" is an artificial intelligence technology for generating avatars based on basic information about users.
[1470] An "avatar" is a digital representation with a character that is generated based on basic information about the user.
[1471] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates dialogue content based on user input.
[1472] "Emotion analysis" is a technology that analyzes emotions from user input and extracts meaning.
[1473] An "empathetic response" is a response that is generated in a way that is in tune with the user's emotions, based on analyzed emotional information.
[1474] The "database" is an information system that stores the generated conversation and emotion information and uses it to update the content of future conversations.
[1475] "Encryption" is a data protection method used to safely store users' personal information.
[1476] The present invention provides a system for providing personalized empathetic interaction to users who feel lonely. The system is implemented according to the following steps.
[1477] Initial Setup and Avatar Creation
[1478] Users access the platform via a web browser or a dedicated app. After accessing the platform, users enter basic information (name, age, hobbies and interests) into the displayed input form. This information is then sent from the device to the server.
[1479] The server creates a user profile based on the received basic information. Using this profile information, it activates an image generation AI (e.g., Stable Diffusion) to generate an avatar corresponding to the user. The generated avatar is sent to the terminal and displayed to the user. The user receives a display to confirm this avatar.
[1480] Conversation initiation and emotional understanding
[1481] When the user clicks the "Start" button, the operation is notified from the terminal to the server.
[1482] The server launches a conversational AI (e.g., GPT-4) to generate an initial dialogue message based on the user's basic information. For example, a message like "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" is generated. The device displays the message to the user. When the user responds with text or voice, that information is sent from the device to the server.
[1483] The server analyzes the input information using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). Through this analysis, the server identifies the user's emotions (e.g., interest, joy). Based on the analysis results, the server generates an appropriate empathetic response. For example, if the user responds, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. What are your favorite mystery novels?" This response is displayed to the user via their device.
[1484] Continuous personalization
[1485] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This generates basic data to further personalize the content of the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information is stored in the database and reflected in future conversations. In the next conversation, a personalized message such as, "Regarding the painting we talked about last time, have you painted any recent works?" is generated.
[1486] Safety and Privacy Controls
[1487] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users have the right to view, correct, or delete their own data within the platform, and the server updates the information in the database accordingly.
[1488] Examples and prompts
[1489] Example 1: First conversation
[1490] 1. The user enters "Yamada Taro, 30 years old, hobby is reading."
[1491] 2. The server creates a profile for "Yamada Taro" and activates an image-generating AI (e.g., Stable Diffusion) to generate an avatar.
[1492] 3. The device displays the avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1493] Example prompt sentence:
[1494] Prompt for generative AI models (image-generating AI):
[1495] "Please create an avatar for a 30-year-old Japanese man whose hobby is reading, with a kind and intelligent face."
[1496] Prompts for generative AI models (generative conversational AI):
[1497] "The user's name is Taro Yamada, and his hobby is reading. As an initial greeting to him, generate a message that says, 'Hello, Taro Yamada. I heard you like reading. What book have you read recently?'"
[1498] Example 2: Continuation Conversation
[1499] 1. The user types, "I've been reading mystery novels lately."
[1500] 2. The server analyzes the input content and determines it as "interest" using the emotion recognition module.
[1501] 3. The server generates an empathetic response such as, "That's interesting. Which mystery novels did you particularly enjoy?" and displays it on the device.
[1502] Example prompt sentence:
[1503] Prompt to the generative AI model (sentiment analysis module):
[1504] "Text: I recently read a mystery novel. Please analyze my emotions and tell me what would be an appropriate response."
[1505] Prompts for generative AI models (generative conversational AI):
[1506] "A user types, 'I've been reading mystery novels lately.' Generate an empathetic response like, 'That sounds interesting. Which mystery novels did you enjoy the most?'"
[1507] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[1508] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1509] Step 1:
[1510] A user accesses the platform and enters basic information (name, age, hobbies and interests). The basic information entered by the user is sent from the device to the server. The input data includes the user's name (e.g., Yamada Taro), age (e.g., 30 years old), hobbies (e.g., reading), etc. This data is used to create the user's profile.
[1511] Input: User's basic information (name, age, hobbies and interests)
[1512] Output: Basic information sent to the server
[1513] Step 2:
[1514] The server creates a user profile based on the received basic information. This includes storing the basic information in a database. It then invokes an image-generating AI (e.g., Stable Diffusion) to generate an avatar for the user. The generated avatar has characteristics based on the basic information, such as an intelligent and kind appearance. The generated avatar is sent from the server to the device and displayed to the user. It is displayed along with the message, "Are you sure you want this avatar?"
[1515] Input: Basic information sent to the server
[1516] Output: An avatar image and a confirmation message that will be displayed to the user.
[1517] Step 3:
[1518] When the user clicks the "Start" button, the device notifies the server of the operation. The server receives this notification and launches a conversation-generating AI (e.g., GPT-4). The server generates an initial greeting and question based on the user's basic information. For example, it generates a message like, "Hello, Yamada-san. I heard you like reading. What book have you read recently?" This message is sent from the server to the device and displayed to the user.
[1519] Input: User interaction start operation
[1520] Output: Initial greeting message and question
[1521] Step 4:
[1522] When the user responds with text or voice, the input is sent from the device to the server. The server analyzes the input using an emotion analysis module (e.g., IBM Watson Natural Language Understanding). The analysis identifies the user's emotion (e.g., interest, joy). The extracted emotion information is used in the next step.
[1523] Input: User response (text or voice)
[1524] Output: Parsed emotion data
[1525] Step 5:
[1526] The server generates an appropriate empathetic response based on the results of the emotion analysis. It uses conversation-generating AI to create an empathetic message. For example, a response might be generated such as, "That sounds interesting. What are your favorite mystery novels?" This response is sent from the server to the device and displayed to the user.
[1527] Input: Parsed emotion data
[1528] Output: Empathetic response message
[1529] Step 6:
[1530] As the conversation with the user continues, the server stores the conversation history and emotional information in a database. This stored data is used to personalize the next conversation. For example, if the user brings up a new hobby (e.g., painting), that information will be reflected in the next conversation. In the next conversation, a personalized message will be generated, such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[1531] Input: Conversation history and emotional information
[1532] Output: Database information used for future personalization
[1533] Step 7:
[1534] Users' personal information is encrypted (e.g., AES-256) by the server and stored securely. The device displays the necessary information only upon the user's request, and hides all other data. Users can view, modify, and delete their own data within the platform, and the server updates the information in the database accordingly.
[1535] Input: User's personal information
[1536] Output: Encrypted database information and the ability to view, modify, and delete data according to user requests
[1537] (Application example 1)
[1538] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1539] In modern society, the number of users who feel lonely is increasing. To improve customer experience, particularly in brick-and-mortar stores, personalized, empathetic dialogue is needed, rather than simply providing guidance. However, current in-store guidance systems and guide robots have difficulty responding flexibly based on the customer's basic information and emotions. Therefore, the present invention aims to provide a system that reduces customers' feelings of loneliness and improves customer experience by providing personalized, empathetic dialogue.
[1540] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1541] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next conversation based on the stored information, means for providing personalized guidance through continuous dialogue based on the user's basic information, means for implementing the system in a store using a robot, and means for encrypting the user's personal information and securely storing the avatar and its data. This makes it possible to provide personalized empathetic dialogue in a store based on the user's basic information, hobbies, and interests.
[1542] "Basic information" refers to the initial data needed to individualize and personalize the interaction, such as the user's name, age, hobbies, and interests.
[1543] "Image generation artificial intelligence" is an artificial intelligence technology for generating an avatar corresponding to a user based on basic information about the user.
[1544] An "avatar" is a digital person or character with a virtual appearance and personality that is generated based on basic information about the user.
[1545] "Conversational generative artificial intelligence" is an artificial intelligence technology that generates appropriate dialogue content based on input information from the user.
[1546] "Emotion analysis" is a technology for extracting emotions from user input and determining emotions such as interest, joy, sadness, etc.
[1547] An "empathetic response" is a response that is generated based on the results of analyzing the user's emotions and is given in a manner that empathizes with the user's emotions.
[1548] The "database" is an information management system that stores the generated conversation and emotion information and updates or references it as needed.
[1549] "Encryption" is the technology of converting data into cryptographic code to protect users' personal information.
[1550] "Individualized guidance" refers to guidance information tailored to the individual needs and interests of a user, provided based on the user's basic information and past interaction history.
[1551] A "robot" is a mechanical device that exists physically and has a guidance function to operate the above system in the real world.
[1552] To implement the present invention, the following system configuration and processing procedures are basically required.
[1553] A user first accesses the system and enters their basic information, including name, age, hobbies, and interests. The system then creates a user profile and uses image generation AI to generate an avatar for the user. The terminal then displays the generated avatar to the user and asks for their confirmation.
[1554] When the user clicks the "Start" button to begin a conversation with the avatar, the device notifies the server of this action. The server then activates a conversation-generating AI to generate an initial greeting and question based on the user's basic information. For example, a message such as "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" may be generated.
[1555] When the user responds with text or voice, the device sends the input to the server. The server analyzes the input using an emotion analysis module to extract the user's emotion. An empathetic response is generated based on the analyzed emotion. For example, if the user inputs, "I've been reading mystery novels lately," the server generates a response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?" This response is displayed to the user via the device.
[1556] This system realizes continuous personalized dialogue by storing the content of the dialogue with the user and emotional information in a database. The next time the dialogue is held, the content is updated based on the stored data, providing a more personalized response. For example, if the user mentioned a new hobby (e.g., painting) in the previous dialogue, the next dialogue could continue with a question such as, "Regarding the painting we talked about last time, have you painted any works recently?"
[1557] The system encrypts users' personal information and stores avatars and related data securely, protecting users' privacy. Users can also view, modify, and delete their own data within the platform.
[1558] As a concrete example, the following system usage scenarios can be considered:
[1559] Example 1: When a guide robot approaches a customer in a store
[1560] "Hello, Mr. / Ms. XX. Thank you for coming. I heard you enjoy reading. What book have you read recently?"
[1561] An example of a prompt is:
[1562] User: Yamada Taro. Age: 30. Hobbies: reading and watching movies.
[1563] What mystery novel have you read recently?”
[1564] In this way, the system can provide users with a personalized and empathetic interaction experience, reducing feelings of loneliness.
[1565] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1566] Step 1:
[1567] A user accesses the system and enters their basic information. This basic information includes name, age, hobbies, interests, etc. The entered basic information is sent from the terminal to the server. Input: User's basic information. Output: Basic information saved on the server.
[1568] Step 2:
[1569] The server creates a user profile based on the received basic information. An image generation AI uses this user profile to generate an avatar. The generated avatar is sent to the device. Input: Basic user information. Output: Generated avatar.
[1570] Step 3:
[1571] The device displays the generated avatar to the user and asks for confirmation. Input: Generated avatar. Output: Displayed avatar and user confirmation.
[1572] Step 4:
[1573] When the user clicks the "Start" button to start interacting with the avatar, the terminal notifies the server of this operation. Input: User operation. Output: Notification to the server.
[1574] Step 5:
[1575] The server starts a conversation generation AI to generate an initial greeting and question based on the user's basic information. The generated message is sent to the terminal and displayed to the user. Input: User's basic information. Output: Generated initial message.
[1576] Step 6:
[1577] The user responds with text or voice. This input is sent from the device to the server. Input: User response. Output: Response sent to the server.
[1578] Step 7:
[1579] The server analyzes the input content with an emotion analysis module to extract the user's emotion. The analysis result is used as input for generating an empathetic response. Input: User's response. Output: Analyzed emotion.
[1580] Step 8:
[1581] The server generates an empathetic response based on the emotion analysis results and sends the response to the device. The device displays the generated response to the user. Input: Analyzed emotion. Output: Generated empathetic response.
[1582] Step 9:
[1583] The server saves the generated conversation and emotion information in a database. The next time the conversation is held, the content of the conversation will be updated based on this information. Input: Generated conversation and emotion information. Output: Information saved in the database.
[1584] Step 10:
[1585] The server encrypts all data and stores it securely. Users can view, modify, or delete their own data at any time. Input: Information stored in the database. Output: Encrypted data.
[1586] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1587] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely, and by combining an emotion engine, recognizes the user's emotions and generates responses accordingly. The program and processing flow of this system are described in detail below.
[1588] Initial Setup and Avatar Creation
[1589] First, the user accesses the platform and enters basic information (such as name, age, hobbies and interests). The device then sends this information to the server. The server then creates a user profile based on the received basic information and calls on image-generating AI to generate an avatar. The avatar is designed to have the optimal appearance and character based on the information entered by the user. The device then displays the generated avatar to the user and asks for confirmation.
[1590] Conversation initiation and emotion recognition
[1591] When a user starts their first conversation with an avatar, they click the "Start" button. This action is notified to the server from the device, and the server activates a conversation generation AI to generate an initial greeting and question based on the user's basic information. For example, it might generate a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?" In response, the user inputs text or voice, and the device sends the input to the server.
[1592] Emotion recognition by emotion engine
[1593] The server analyzes the received input content using an emotion engine to identify the user's emotion (e.g., interest, joy, sadness). The emotion engine extracts emotional signals from various inputs such as text, voice, and facial expressions to determine the user's emotional state. For example, if a user inputs, "I've been reading mystery novels recently," the emotion engine will identify "interest" and "joy."
[1594] Generating empathetic responses and interacting
[1595] Based on this analyzed emotional information, the server generates an empathetic response, such as, "That sounds interesting. Which mystery novel did you particularly enjoy?" Furthermore, the avatar's facial expression and attitude change according to this emotional information, visually demonstrating empathy.
[1596] The generated response message is displayed to the user via the terminal, and this dialogue cycle is repeated.Furthermore, the server stores the generated conversation and emotion information in a database and manages the conversation history.
[1597] Continuous personalization
[1598] The server optimizes the content of the next conversation based on the history in the database. It learns changes in the user's hobbies and interests and reflects these in the content of the conversation. For example, if the user talks about a new hobby (e.g., painting), the server continues the conversation by asking, "Regarding the painting we talked about last time, have you painted any works recently?"
[1599] Safety and Privacy Controls
[1600] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon the user's request, and hides other data. Users can view their own data within the platform and modify or delete it as necessary. The server updates the information in the database in response to this request, protecting privacy.
[1601] Specific examples
[1602] Example 1: First conversation
[1603] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1604] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1605] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1606] Example 2: Continuation Conversation
[1607] A user types, "I've been reading mystery novels lately."
[1608] The server analyzes the input using an emotion engine and determines it as "interest."
[1609] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1610] The terminal displays this message to the user.
[1611] In this way, the system provides a personalized and empathetic interaction experience for the user, reducing feelings of loneliness and leveraging the emotion engine to further address the user's emotional state.
[1612] The processing flow will be explained below.
[1613] Step 1:
[1614] The user accesses the platform and enters basic information (such as name, age, hobbies and interests), which is then sent to the server.
[1615] Step 2:
[1616] The server creates a user profile based on the received basic information, then invokes an image generation AI to generate an avatar based on the input information.
[1617] Step 3:
[1618] The device displays the generated avatar to the user for confirmation, and if the user is not satisfied with the displayed avatar, they can request corrections or regeneration.
[1619] Step 4:
[1620] The user clicks the "Start" button to begin interacting with the avatar. This action is notified to the server from the device.
[1621] Step 5:
[1622] The server then activates a conversation-generating AI to generate an initial greeting and questions based on the user's basic information. For example, it generates a message like, "Hello, Mr. / Ms. XX. I heard you like reading. What books have you read recently?"
[1623] Step 6:
[1624] The device displays the generated message to the user, who can then respond with text or voice input.
[1625] Step 7:
[1626] The device sends the user's input (text or voice) to the server.
[1627] Step 8:
[1628] The server analyzes the received input using an emotion engine to identify the user's emotions, such as "joy," "interest," and "sadness."
[1629] Step 9:
[1630] The server then sends commands to the device to change the avatar's facial expression and attitude based on the extracted emotional information. For example, if the user inputs "I've been reading mystery novels recently," the avatar will display expressions and attitudes that express "interest" or "joy."
[1631] Step 10:
[1632] The server generates an empathetic response based on the emotional information, for example, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1633] Step 11:
[1634] The device then displays the generated empathetic response message to the user, and this cycle is repeated for each interaction with the user.
[1635] Step 12:
[1636] The server stores the content of the conversation with the user and emotional information in a database and manages the conversation history.
[1637] Step 13:
[1638] The server extracts information from the history in the database to optimize the content of the next conversation, updating it to reflect the user's new hobbies and interests.
[1639] Step 14:
[1640] Users can view their own data within the platform and modify or delete it as necessary, and the device will notify the server of this action.
[1641] Step 15:
[1642] The server updates the information in the database according to user requests and also manages encrypted data for security purposes.
[1643] Specific examples
[1644] Example 1: First conversation
[1645] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1646] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1647] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1648] Example 2: Continuation Conversation
[1649] A user types, "I've been reading mystery novels lately."
[1650] The server analyzes the input using an emotion engine and determines it as "interest."
[1651] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1652] The device displays a response message, and the avatar's facial expression and attitude change to reflect interest.
[1653] Example 2
[1654] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1655] In modern society, the number of people feeling lonely is increasing, and there is a need for systems that provide empathetic dialogue tailored to individual needs. However, conventional dialogue systems have difficulty fully analyzing users' emotions and providing empathetic responses. They also fall short in terms of protecting the safety and privacy of users' personal information. A system that solves these problems and reduces feelings of loneliness is needed.
[1656] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1657] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated dialogue and emotional information in a database, means for updating the content of the next dialogue based on the stored information, means for encrypting the user's personal information and safely storing the data, means for changing the avatar's facial expression and attitude based on the emotional information, means for displaying necessary information and hiding other data in response to a user request, and means for the user to view, modify, and delete data. This makes it possible to provide an empathetic dialogue that is responsive to the user's emotions, while protecting the safety and privacy of the personal information.
[1658] "User" refers to an individual who interacts with the system.
[1659] "Basic Information" refers to personal information entered by the user, such as name, age, hobbies and interests.
[1660] A "terminal" is a device that is operated by a user to input and display information.
[1661] "Server" refers to a central computer system that receives and processes information sent by users.
[1662] "Image generation artificial intelligence" refers to artificial intelligence technology for generating avatars based on basic user information.
[1663] An "avatar" is a virtual persona or character created based on a user's basic information.
[1664] "Conversational generative artificial intelligence" refers to artificial intelligence that has the technology to generate dialogue content based on user input.
[1665] An "emotion engine" refers to algorithms and technologies for analyzing and identifying emotions from user input.
[1666] "Emotion information" refers to data on the user's emotional state analyzed by the emotion engine.
[1667] "Empathetic response" refers to a dialogue that is based on analyzed emotional information and is sensitive to the user's emotions.
[1668] "Database" refers to a data structure for storing and managing generated conversation and emotion information.
[1669] "Encryption" refers to the technology of converting digital data into a form that cannot be read by third parties.
[1670] "Safe storage" refers to measures to protect users' personal information and data from unauthorized access and information leaks.
[1671] "Changing facial expressions and attitudes" refers to changing the visual response of the avatar displayed to the user based on emotional information.
[1672] "Displaying information on demand" refers to displaying only the necessary information based on the user's instructions.
[1673] "Viewing, modifying, and deleting data" refers to the actions that users can take to view, modify, or delete their own data through the platform.
[1674] The present invention is a system that provides personalized empathetic dialogue to users who feel lonely. This system has the following configuration to recognize the user's emotions by combining emotion engines and generate responses according to those emotions.
[1675] Initial Setup and Avatar Creation
[1676] First, a user accesses the platform and enters basic information such as name, age, hobbies, and interests. The device then sends this information to the server. The server creates a user profile based on the received information and calls on image-generating AI (e.g., DeepArt) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input information.
[1677] The generated avatar is displayed to the user via the terminal and they are asked to confirm. Through this process, the user can obtain an avatar that suits their preferences.
[1678] Conversation initiation and emotion recognition
[1679] To begin a conversation with an avatar, a user clicks the "Start" button. This action causes the device to notify the server. The server then launches a conversation-generating AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[1680] For example, the server might generate a message like, "Hello, [username]. I heard you like reading. What book have you read recently?" The user can respond with text or voice, and the device will send that input to the server.
[1681] Emotion recognition by emotion engine
[1682] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer), which extracts emotional signals from inputs such as text, voice, and facial expressions to identify the user's emotions (interest, joy, sadness, etc.).
[1683] For example, if a user types, "I recently read a mystery novel," the emotion engine identifies emotions like "interest" and "joy," and based on this, the server generates a more empathetic response.
[1684] Generating empathetic responses and interacting
[1685] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds interesting. Which mystery novel did you particularly enjoy?" The avatar's facial expressions and attitude also change according to this emotional information, visually demonstrating empathy.
[1686] The generated responses are displayed to the user via the device, and this dialogue cycle is repeated. At the same time, the server stores the generated conversation and emotional information in a database and manages the conversation history. This history is used to optimize the content of the next dialogue.
[1687] Continuous personalization
[1688] The server optimizes the content of the next conversation based on the saved history data. This allows the server to learn about changes in the user's hobbies and interests and reflect that information in the conversation. For example, if the user takes up a new hobby (e.g., painting), the conversation can continue with a prompt such as, "Regarding the painting we talked about last time, have you painted any recent works?"
[1689] Safety and Privacy Controls
[1690] Users' personal information is encrypted and securely stored by the server. The device displays necessary information only upon user request, hiding other data. Furthermore, users can view their own data within the platform and modify or delete it as needed. The server updates the information in the database in response to these requests, protecting privacy.
[1691] Specific examples
[1692] Example 1: First conversation
[1693] The user enters "Yamada Taro, 30 years old, hobby is reading."
[1694] The server creates a profile for "Yamada Taro" and activates image-generating artificial intelligence to generate an avatar.
[1695] The device displays an avatar and a server-generated message: "Hello, Yamada-san. I heard you like reading. What book have you read recently?"
[1696] Example 2: Continuation Conversation
[1697] A user types, "I've been reading mystery novels lately."
[1698] The server analyzes the input using an emotion engine and determines it as "interest."
[1699] The server generates an empathetic response such as, "That sounds interesting. Which mystery novels did you particularly enjoy?"
[1700] The terminal displays this message to the user.
[1701] In this way, the system of the present invention can provide a personalized and empathetic interaction experience for the user, reducing feelings of loneliness. Furthermore, by utilizing the emotion engine, the system can generate appropriate responses to the user's emotional state, improving the quality of the interaction.
[1702] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1703] Step 1: Enter basic information
[1704] Users access the platform and enter basic information such as their name, age, hobbies and interests.
[1705] Input: Basic information such as name, age, hobbies and interests
[1706] Output: Basic information data
[1707] Step 2: Submit basic information
[1708] The terminal transmits the input basic information to the server.
[1709] Input: Basic information data
[1710] Output: Basic information data sent to the server
[1711] Step 3: Create your profile and create your avatar
[1712] The server creates a user profile based on the received basic information and calls an image-generating artificial intelligence (e.g., DeepArt) to generate an avatar.
[1713] Input: Basic information data sent to the server
[1714] Output: Generated avatar data
[1715] Step 4: View and check your avatar
[1716] The terminal displays the generated avatar to the user and asks for confirmation.
[1717] Input: Generated avatar data
[1718] Output: The avatar displayed to the user
[1719] Step 5: Initiating an initial conversation
[1720] The user clicks the "Get Started" button.
[1721] The terminal notifies the server of this operation.
[1722] Input: User clicks the "Get Started" button
[1723] Output: Notification to the server
[1724] Step 6: Generate initial greetings and questions
[1725] The server launches a conversational AI (e.g., OpenAI's GPT-4) to generate an initial greeting and questions based on the user's basic information.
[1726] Input: Notification to the server, basic information data of the user
[1727] Output: Initial greeting and question message
[1728] Step 7: User response input
[1729] The user enters a response by text or voice.
[1730] The terminal transmits the input response content to the server.
[1731] Input: User response
[1732] Output: Response sent to the server
[1733] Step 8: Emotion Recognition
[1734] The server analyzes the received input using an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the emotion.
[1735] Input: Response sent to the server
[1736] Output: Identified emotion information
[1737] Step 9: Generate an empathetic response
[1738] The server generates an empathetic response based on the analyzed emotional information, such as "That sounds interesting. Which mystery novel did you particularly enjoy?", and also changes the avatar's facial expression and attitude according to this emotional information.
[1739] Input: Identified emotion information
[1740] Output: Empathetic response message, modified avatar data
[1741] Step 10: View the response
[1742] The generated response message is displayed to the user through the terminal, along with the avatar's facial expression and attitude.
[1743] Input: Empathetic response message, modified avatar data
[1744] Output: Response message displayed to the user and the avatar's facial expression and attitude
[1745] Step 11: Saving conversation history
[1746] The server stores the generated conversation and emotion information in a database.
[1747] Input: Generated conversation data, identified emotion information
[1748] Output: Conversation data and emotion information stored in a database
[1749] Step 12: Optimize the content of your next conversation
[1750] The server optimizes the content of the next conversation based on the stored historical data.
[1751] Input: Conversation data and emotion information stored in a database
[1752] Output: Optimized next dialogue
[1753] Step 13: Managing your personal information
[1754] The server encrypts and securely stores users' personal information.
[1755] The device displays necessary information only when requested by the user, and hides other data. The user can also view, modify, and delete data.
[1756] Input: User's personal information, User's request
[1757] Output: Encrypted personal information, securely stored data, data displayed, modified, or deleted upon user request
[1758] Through these steps, it is possible to provide users with personalized, empathetic interactions and reduce their feelings of loneliness. Furthermore, by utilizing the emotion engine, it is possible to generate appropriate responses according to the user's emotional state, improving the quality of the interactions.
[1759] (Application example 2)
[1760] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1761] Conventional food delivery applications lacked a way to address the feelings of loneliness felt by users when ordering food. Furthermore, they were unable to provide empathetic dialogue that responded to the user's emotional state, making it difficult to improve the quality of the user experience.
[1762] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1763] In this invention, the server includes means for inputting basic information from a user, means for generating an avatar using an image generation type artificial intelligence based on the input basic information, means for displaying the generated avatar to the user, means for activating a conversation generation type artificial intelligence based on the user's input and engaging in a dialogue, means for analyzing emotions from the user's input, means for generating an empathetic response based on the analyzed emotional information, means for storing the generated conversation and emotional information in a database, means for updating the next conversation content based on the stored information, means for encrypting the user's personal information and safely storing the avatar and its data, means for providing empathetic dialogue if the user feels lonely based on the input content when ordering a meal, and means for analyzing the user's emotions when ordering a meal using an emotion engine and suggesting a response based on the emotions. This makes it possible to provide psychological support through empathetic dialogue even when the user feels lonely, thereby improving the quality of the user experience.
[1764] "Basic information" is personal information such as the user's name, age, hobbies and interests.
[1765] "Image generation artificial intelligence" is an artificial intelligence technology that generates images based on input data.
[1766] An "avatar" is a virtual character created to represent a user.
[1767] "Conversational artificial intelligence" is an artificial intelligence technology that generates natural dialogue based on user input.
[1768] An "emotion engine" is a technology for analyzing emotions from user input and behavior.
[1769] An "empathetic response" is a response that understands the user's emotions and is generated in a way that is sympathetic to those emotions.
[1770] The "database" is an information storage system for storing users' conversation history and emotional information.
[1771] "Encryption" is a technology that converts data to safely store users' personal information and prevents external access.
[1772] "Food delivery" is a service that allows users to order meals online and have them delivered to a specified location.
[1773] "Loneliness" is a psychological state in which one feels disconnected from others and isolated.
[1774] "Psychological support" is assistance to provide users with psychological stability and a sense of security.
[1775] This invention is a system for realizing a food delivery application that provides empathetic dialogue when a user feels lonely when ordering a meal. This system operates using a smartphone and a server.
[1776] First, a user accesses the application using their smartphone and enters basic information such as name, age, hobbies and interests, which is then sent from the device to a server, which then creates a user profile.
[1777] Next, the server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar. This avatar is designed to have the optimal appearance and character based on the user's input. The generated avatar is displayed on the device and the user is asked to confirm it.
[1778] When a user starts interacting with an avatar, the server activates a conversational AI (e.g., GPT-3) to generate an initial greeting and question. Based on the user's basic information, the server generates empathetic messages such as "Hello, how busy have you been?" or "Do you have any particular favorite dishes?" The user inputs the information as text or voice, which is then sent from the device to the server.
[1779] The server analyzes the user's input using an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to identify the user's emotion. For example, if a user inputs, "I often eat alone, and I feel a little lonely," the emotion engine will identify the emotion as "loneliness."
[1780] Based on the analyzed emotional information, the server generates an empathetic response, such as "That sounds lonely. What kind of person would be good to be with at a time like this?" The generated response is then displayed to the user via the device.
[1781] This cycle repeats, and the server stores the generated conversation and emotional information in a database, optimizing the content of the next conversation based on the stored information and adapting to changes in the user's hobbies and interests.
[1782] In addition, users' personal information is encrypted and securely stored by the server. This allows users to view, modify, and delete their data within the platform. The device will only display necessary information upon user request, and will hide other data.
[1783] As a concrete example, the following cases can be considered:
[1784] Example 1: First conversation
[1785] The user types, "Hello, I'm Taro Tanaka, I'm 25 years old, and my hobby is cooking."
[1786] The server generates the message "Hello, Tanaka-san. I heard you like cooking. What dish have you made recently?"
[1787] Example 2: Disconnected conversations
[1788] The user types, "I often eat alone and feel a little lonely."
[1789] The server generates an empathetic response: "That sounds lonely. Who do you think would be good to have around at a time like this?"
[1790] Prompt Sentence Examples
[1791] User: "Hi, I'm very tired."
[1792] AI: "I see, that's tough. What was particularly tiring for you today?"
[1793] In this way, the system of the present invention provides users with an empathetic interactive experience to reduce feelings of loneliness and realize psychological support.
[1794] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1795] Step 1:
[1796] The user uses a smartphone to enter basic information (name, age, hobbies, interests, etc.).
[1797] Input: Name, age, hobbies and interests
[1798] Output: Basic information set (user information completed)
[1799] What happens: A user enters the required information into a form and clicks the submit button.
[1800] Step 2:
[1801] The terminal sends the user's basic information to the server.
[1802] Input: Basic information
[1803] Output: Send data to the server
[1804] Operation: The device encrypts the basic information entered and sends it to the server.
[1805] Step 3:
[1806] The server uses image generation artificial intelligence to generate an avatar based on basic information.
[1807] Input: Encrypted basic information
[1808] Output: The generated avatar
[1809] How it works: The server calls an image-generating AI (e.g., DALL-E or GAN) to generate an avatar that reflects the user's characteristics.
[1810] Step 4:
[1811] The terminal displays the generated avatar to the user and asks for confirmation.
[1812] Input: Generated avatar
[1813] Output: Display avatar and confirm user
[1814] What it does: The device displays the avatar on the screen and prompts the user to confirm the avatar.
[1815] Step 5:
[1816] The user clicks the "Start" button to begin interacting with the avatar.
[1817] Input: Start Operation
[1818] Output: Trigger to start a conversation
[1819] Action: The user clicks the "Get Started" button on the application screen.
[1820] Step 6:
[1821] The server activates a conversation-generating artificial intelligence (AI) to generate initial greetings and questions based on the user's basic information.
[1822] Input: Basic information, trigger to start a conversation
[1823] Output: Initial greetings and questions
[1824] How it works: The server uses a generative conversational AI (e.g., GPT-3) to generate messages such as "Hello, how have you been?"
[1825] Step 7:
[1826] The user inputs text or voice, and the device sends the input to the server.
[1827] Input: User text or voice input
[1828] Output: Sending user input to the server
[1829] Operation: The device sends the user's text or voice data to the server.
[1830] Step 8:
[1831] The server analyzes the user's input using an emotion engine and identifies the emotion.
[1832] Input: User text or voice input
[1833] Output: Identified emotion information
[1834] How it works: The server uses an emotion engine (e.g. IBM Watson or Microsoft Azure Emotion API) to analyze emotions from the input.
[1835] Step 9:
[1836] Based on the analyzed emotional information, the server generates an empathetic response.
[1837] Input: Emotion information
[1838] Output: Empathetic response message
[1839] How it works: The server generates an empathetic message based on the emotional information, for example, a response like "That sounds lonely."
[1840] Step 10:
[1841] The generated empathetic response is displayed to the user via the terminal.
[1842] Input: Empathetic response message
[1843] Output: The response message that is displayed to the user
[1844] Action: The terminal displays the response message on the screen.
[1845] Step 11:
[1846] The server stores the generated conversation and emotion information in a database.
[1847] Input: Conversation content, emotional information
[1848] Output: Save to database
[1849] How it works: The server records conversation history and emotion information in a database.
[1850] Step 12:
[1851] The conversation will be updated based on the saved information the next time you interact, and will also respond to changes in the user's hobbies and interests.
[1852] Input: saved conversation history, emotional information
[1853] Output: Optimized dialogue
[1854] How it works: The server retrieves information from the database and personalizes your next interaction.
[1855] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1856] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1857] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1858] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1859] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1860] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1861] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1862] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1863] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1864] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1865] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1866] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1867] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1868] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1869] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1870] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1871] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1872] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1873] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1874] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1875] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1876] The following is further disclosed regarding the above embodiment.
[1877] (Claim 1)
[1878] a means for inputting basic information from a user;
[1879] A means for generating an avatar using an image generating artificial intelligence based on the input basic information;
[1880] means for displaying the generated avatar to a user;
[1881] a means for activating a conversation generation type artificial intelligence based on an input from a user and engaging in a dialogue;
[1882] means for analyzing emotions from user input;
[1883] means for generating an empathetic response based on the analyzed emotional information;
[1884] a means for storing the generated conversation and emotion information in a database;
[1885] A way to update the next conversation based on the saved information, and
[1886] A system that encrypts users' personal information and includes a means to securely store avatars and their data.
[1887] (Claim 2)
[1888] 10. The system of claim 1, further comprising means for learning changes in the user's hobbies and interests and optimizing the content of the dialogue based on the learning changes.
[1889] (Claim 3)
[1890] 10. The system of claim 1, further comprising means for dynamically updating conversation content in response to new hobbies or interests of the user.
[1891] "Example 1"
[1892] (Claim 1)
[1893] a means for inputting basic information from a user;
[1894] A means for generating an avatar using an image generating artificial intelligence based on the input basic information;
[1895] means for displaying the generated avatar to a user;
[1896] a means for activating a conversation generation type artificial intelligence based on an input from a user and engaging in a dialogue;
[1897] means for analyzing emotions from user input;
[1898] means for generating an empathetic response based on the analyzed emotional information;
[1899] a means for storing the generated conversation and emotion information in a database;
[1900] A way to update the next conversation based on the saved information, and
[1901] A means to encrypt users' personal information and store avatars and their data securely;
[1902] Terminal means through which users access the platform and enter basic information;
[1903] means for transmitting the entered information to a server, which then creates a profile;
[1904] a means for displaying the generated avatar to the user on the terminal and requesting confirmation;
[1905] A means for the terminal to notify the server of a user's interaction start operation;
[1906] means for transmitting user responses during a conversation from the terminal to the server;
[1907] The system includes a means for generating an appropriate empathetic response based on the emotion analysis results.
[1908] (Claim 2)
[1909] 10. The system of claim 1, further comprising means for learning changes in the user's hobbies and interests and optimizing the content of the dialogue based on the learning changes.
[1910] (Claim 3)
[1911] 10. The system of claim 1, further comprising means for dynamically updating conversation content in response to new hobbies or interests of the user.
[1912] "Application Example 1"
[1913] (Claim 1)
[1914] a means for inputting basic information from a user;
[1915] A means for generating an avatar using an image generating artificial intelligence based on the input basic information;
[1916] means for displaying the generated avatar to a user;
[1917] a means for activating a conversation generation type artificial intelligence based on an input from a user and engaging in a dialogue;
[1918] means for analyzing emotions from user input;
[1919] means for generating an empathetic response based on the analyzed emotional information;
[1920] a means for storing the generated conversation and emotion information in a database;
[1921] A way to update the next conversation based on the saved information, and
[1922] A means for providing personalized guidance through continuous dialogue based on basic user information;
[1923] means for implementing the system in a store using a robot;
[1924] A system that encrypts users' personal information and includes a means to securely store avatars and their data.
[1925] (Claim 2)
[1926] 10. The system of claim 1, further comprising means for learning changes in the user's hobbies and interests and optimizing the content of the dialogue based on the learning changes.
[1927] (Claim 3)
[1928] 10. The system of claim 1, further comprising means for dynamically updating conversation content in response to new hobbies or interests of the user.
[1929] "Example 2: Combining Emotion Engines"
[1930] (Claim 1)
[1931] a means for inputting basic information from a user;
[1932] A means for generating an avatar using an image generating artificial intelligence based on the input basic information;
[1933] means for displaying the generated avatar to a user;
[1934] a means for activating a conversation generation type artificial intelligence based on an input from a user and engaging in a dialogue;
[1935] means for analyzing emotions from user input;
[1936] means for generating an empathetic response based on the analyzed emotional information;
[1937] a means for storing the generated conversation and emotion information in a database;
[1938] A way to update the next conversation based on the saved information, and
[1939] A means of encrypting your personal information an...
Claims
1. a means for inputting basic information from a user; A means for generating an avatar using an image generating artificial intelligence based on the input basic information; means for displaying the generated avatar to a user; a means for activating a conversation generation type artificial intelligence based on an input from a user and engaging in a dialogue; means for analyzing emotions from user input; means for generating an empathetic response based on the analyzed emotional information; a means for storing the generated conversation and emotion information in a database; A way to update the next conversation based on the saved information, and A system that encrypts users' personal information and includes a means to securely store avatars and their data.
2. 2. The system of claim 1, further comprising means for learning changes in the user's hobbies and interests and optimizing the content of the dialogue based thereon.
3. 10. The system of claim 1, further comprising means for dynamically updating conversation content in response to new hobbies or interests of the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A