System
The system addresses loneliness and isolation by enabling virtual conversations with generative AI models that simulate realistic dialogues based on user-input information, providing emotional healing and a sense of security.
Patent Information
- Application Number
- JP2024128500
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
Modern society faces issues of loneliness and isolation due to declining birthrates, aging populations, and changes in living environments, with limited means to obtain emotional healing and psychological support, especially in situations where reunions with estranged friends or loved ones are not possible.
A system that enables virtual conversations with estranged friends or loved ones using generative AI models to simulate realistic dialogues, allowing users to input information about the conversation target, which is stored in a database and used to train an AI model for natural language analysis and response generation.
Provides emotional healing and a sense of security through realistic and emotionally charged virtual conversations that reflect the personality and past memories of the conversation target, alleviating feelings of loneliness and isolation.
Smart Images

Figure 2026025688000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the declining birthrate, aging population, and changes in people's living environments are causing an increasing number of people to feel isolated and lonely. In particular, in situations where it is difficult to reunite with estranged friends or loved ones who have passed away, there are limited means to obtain emotional healing and psychological support. This problem can have a serious impact on mental health and quality of life. The problem that this invention aims to solve is to alleviate these feelings of loneliness and isolation through virtual conversations and provide users with emotional healing. [Means for solving the problem]
[0005] To solve the above-mentioned problems, the present invention provides the following means. A system is provided that includes: a means for inputting information about a conversation target; a means for transmitting the information to a server; a means for the server to store the information in a database; a means for reading the target information from the database and training an AI model; a means for initiating a chat session and having a dialogue between the user and the AI model; and a means for displaying the content of the dialogue to the user. This system allows users to enjoy virtual conversations with estranged friends or loved ones they cannot meet, providing emotional healing and a sense of security. Furthermore, the AI model generates dialogue patterns that reflect the conversation target's personality and past memories, and performs natural language analysis of messages from the user to generate appropriate responses, thereby achieving a natural and friendly dialogue experience.
[0006] "Information about the person you are talking to" refers to basic information about the person you set as your virtual conversation partner, specifically data including name, gender, age, personality, past memories, etc.
[0007] A "server" is a computer system that receives information sent by users and uses that information to train an AI model.
[0008] A "database" is an information management system where a server stores information about the person being spoken to and later provides the data for the AI model to learn from.
[0009] An "AI model" is an artificial intelligence system that learns based on specified conversation target information and enables natural dialogue with the user.
[0010] A "chat session" is a time period during which a user and an AI model engage in a virtual conversation, including the entire dialogue from start to finish.
[0011] A "user" is a person who uses the system to engage in a virtual dialogue with a target person, inputting information through a terminal and communicating with the AI model.
[0012] A "terminal" is a device through which a user inputs subject information, transmits the information to a server, and receives and displays the response generated by the AI model.
[0013] A "dialogue pattern" is a set of response rules or templates that allow the dialogue with the user to proceed naturally, based on the personality and past memories of the person the AI model has learned.
[0014] "Natural language analysis" is a technology in which an AI model analyzes the structure and meaning of language to understand messages entered by users and generate appropriate responses. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system uses generative AI to simulate realistic conversations and is implemented according to the following steps:
[0037] Overall overview
[0038] First, a user enters information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, past memories, etc. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and prepares it for a virtual conversation with the user. When a user starts a chat, a message from the device is sent to the server, and the server uses the AI model to generate an appropriate response, which is then sent back to the user's device. In this way, the virtual conversation progresses in real time.
[0039] Program processing
[0040] User:
[0041] 1. Enter information about the person you are talking to. For example, enter the name, gender, age, etc. of your college friend Ryo, and write in detail about your memories with him (e.g., a college trip).
[0042] 2. Check the information you entered and click the send button.
[0043] Device:
[0044] 1. The entered target information is sent to the server. The data is formatted in JSON and sent securely via HTTPS.
[0045] server:
[0046] 1. Store the received information in a database. Check the data integrity, for example, to make sure that name, age, and gender are entered correctly.
[0047] 2. Based on the saved information, the AI model is activated. The target person information is read from the database and provided as input data to the AI model.
[0048] 3. Based on the stored information, the AI model learns the subject's personality and past memories.
[0049] User:
[0050] 1. Open the chat app and tap the button to start a conversation with the person you set (e.g., "Ryo").
[0051] Device:
[0052] 1. Send a chat session start request to the server. The request includes the user ID and target ID.
[0053] server:
[0054] 1. Start a chat session by generating a new session ID, launching the trained AI model, and associating it with the session ID.
[0055] 2. Receive a message from the user and have the AI model process it. For example, if a user types "How are you doing lately?", the AI model will analyze this message and generate an appropriate response.
[0056] 3. A response is generated and sent to the device. For example, a response such as "It's been a while! Work has been busy lately..." is generated.
[0057] Device:
[0058] 1. Receive a response from the server and display it on the chat screen.
[0059] Specific examples
[0060] Example 1: Conversation with a friend
[0061] When a user wants to have a virtual conversation with a university friend named Ryo, they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, which can then recreate natural conversations such as casual conversations and reminiscing. As the chat progresses, the server generates responses that are characteristic of Ryo in response to questions from the user, providing an experience that feels as if you are having a real conversation.
[0062] Example 2: Conversations with a deceased parent
[0063] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. When the user sends a message such as "Mom, do you remember me?", the server generates a response that reflects the mother's personality and past conversations, such as "Of course! I remember you. We often cooked together when you were little."
[0064] In this way, the present invention can provide a virtual interactive experience that allows users to reflect on past relationships and gain emotional healing and comfort.
[0065] The processing flow will be explained below.
[0066] Understood. The processing steps are explained in detail below.
[0067] Step 1:
[0068] User:
[0069] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[0070] Check the information you entered and click the "Submit" button.
[0071] Step 2:
[0072] Device:
[0073] Format the conversation target information entered by the user into JSON format.
[0074] Information is sent securely to the server using the HTTPS protocol.
[0075] Step 3:
[0076] server:
[0077] The received information is stored in a database.
[0078] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[0079] Step 4:
[0080] server:
[0081] Read the subject's information from the database.
[0082] The read information is provided as input to the AI model, and the AI model is started.
[0083] Step 5:
[0084] server:
[0085] The AI model learns the subject's personality and past memories.
[0086] Generate patterns for dialogue and prepare natural responses.
[0087] Step 6:
[0088] User:
[0089] Open the chat app and tap the "Start Conversation" button.
[0090] Step 7:
[0091] Device:
[0092] A chat session start request is sent to the server, including the user ID and target ID.
[0093] Step 8:
[0094] server:
[0095] Generates a new session ID and initializes the current chat session.
[0096] Associate the trained AI model with a session ID.
[0097] Step 9:
[0098] User:
[0099] Type your first message on the chat screen (e.g., "How are things going lately?").
[0100] Step 10:
[0101] Device:
[0102] Sends the user's message to the server. The message data is sent along with the session ID.
[0103] Step 11:
[0104] server:
[0105] Receive messages from users.
[0106] The message is subjected to natural language analysis, and an AI model generates an appropriate response based on the analysis results.
[0107] Step 12:
[0108] server:
[0109] The response generated by the AI model is sent to the user device (e.g., "It's been a while! Work has been busy lately...").
[0110] Step 13:
[0111] Device:
[0112] Receives a response from the server and displays it on the chat screen.
[0113] Step 14:
[0114] User:
[0115] Continue chatting by continually typing messages (e.g., "Have you traveled anywhere recently?").
[0116] Step 15:
[0117] Device:
[0118] It continues to send each of the user's messages to the server.
[0119] Step 16:
[0120] server:
[0121] Each message is received, natural language analysis is performed, and an appropriate response is generated by the AI model (e.g., "I went to Hokkaido last week. It was so much fun!").
[0122] Step 17:
[0123] Device:
[0124] Each response from the server is received and displayed to the user.
[0125] Step 18:
[0126] User:
[0127] Tap the "End Session" button to end the chat session.
[0128] Step 19:
[0129] Device:
[0130] Sends a request to the server to end the chat session.
[0131] Step 20:
[0132] server:
[0133] Ends the current chat session based on the session ID.
[0134] Archive session data to a database or delete it as needed.
[0135] This allows users to look back on past relationships and find emotional healing through virtual conversations.
[0136] Example 1
[0137] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0138] Loneliness and isolation are serious problems in modern society. In particular, when relationships with close friends and family become distant, or when communication with loved ones who have already passed away is lost, there are limited ways to fill the void. In such situations, users need sufficient support to reflect on past relationships and find emotional healing and security. However, current systems have difficulty providing a more realistic and emotionally charged conversational experience.
[0139] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0140] In this invention, the server includes a means for storing information about the conversation target in a database, a means for reading the conversation target information from the database and training the generative AI model, and a means for starting a new chat session in which the user and the generative AI model converse. This makes it possible to generate conversation patterns that reflect the conversation target's personality and past memories, providing the user with a realistic and emotional conversation experience.
[0141] "Information about the person you are talking to" refers to detailed information about the person you are talking to, such as their name, gender, age, personality, and past memories, entered by the user.
[0142] "Server" refers to the computer system that stores information received from users in a database, trains the generative AI model, and manages chat sessions.
[0143] "Database" refers to a storage system that stores conversation target information and user information for later reference and use.
[0144] A "generative AI model" refers to an artificial intelligence model that learns based on a specific algorithm and generates natural-looking dialogue between a user and a conversation target.
[0145] A "chat session" refers to a series of interactions in which a user initiates a dialogue with a conversation partner and continuously exchanges messages.
[0146] "Session ID" means a unique identifier generated to identify a particular chat session.
[0147] "Natural language analysis" refers to language processing technology for understanding messages from users and generating appropriate responses.
[0148] "Terminal" refers to the device (e.g., smartphone or PC) that a user uses to enter information and communicate with a server.
[0149] "Response" refers to the message generated by the generative AI model and sent to the user.
[0150] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system utilizes generative AI models to simulate realistic conversations and includes the following key elements:
[0151] Hardware and Software Configuration
[0152] Device:
[0153] A device used by a user to input information and communicate with a server. Specifically, this applies to smartphones and PCs.
[0154] server:
[0155] It is a computer system that processes information received from users and generates responses using generative AI models. Its main functions include linking with databases, managing generative AI models, and natural language analysis.
[0156] Database:
[0157] It is a storage system for saving conversation target information and user information for later reference and use.
[0158] Generative AI models:
[0159] It is an artificial intelligence model that learns based on a specific algorithm (e.g., GPT-4) and generates natural dialogue between the user and the person being spoken to.
[0160] System Operation
[0161] user:
[0162] First, a user accesses a chat application using a terminal. Next, they input information about the person they are talking to. This information includes the person's name, gender, age, personality, and past memories. For example, a user can input memories of a friend named "Ryo," a 25-year-old man with a "cheerful and sociable" personality.
[0163] Device:
[0164] Once the user enters the information and clicks the submit button, the device formats this information into JSON format and sends it securely to the server using the HTTPS protocol.
[0165] server:
[0166] The server receives the entered information and stores it in a database. When saving, it checks the integrity of the information, such as name, age, and gender, to ensure it is entered correctly. Once the check is complete, the information is saved in the database.
[0167] The server then reads the target person's information from the database and provides it as input to the generative AI model, which then learns from their personality and past memories to prepare for a conversation with the user.
[0168] When a user starts a chat, the server generates a new session ID, activates the trained generative AI model, and associates it with the session ID. When a message from the user is sent from the device to the server, the server uses the generative AI model to generate an appropriate response and sends it back to the device. The generated response is displayed on the device's chat screen.
[0169] Specific examples
[0170] Scenario of a conversation with a friend
[0171] To enjoy a virtual conversation with a friend from college named "Ryo," a user inputs information about Ryo. For example, the user might enter his name as "Lee," his gender as "male," his age as "25," and mention a "college trip" as a memory. The server uses this information to train a generative AI model. When the user sends a message in a chat app asking "How are you doing lately?", the server generates a response such as "It's been a while! I've been busy at work lately..." and sends it back to the user.
[0172] Scenarios for conversations with deceased parents
[0173] If a user wishes to have a virtual conversation with their deceased mother, they enter detailed information about her (e.g., name, gender, age, personality, memories). The server uses this information to train a generative AI model. When the user sends a message saying, "Mom, do you remember me?", the server generates a response saying, "Of course! I remember you. We used to cook together a lot when you were little," and sends it back to the user.
[0174] The system allows users to reflect on past relationships and find healing and comfort. Using generative AI models, the dialogue progresses in real time, providing users with a highly natural and emotionally rich conversational experience.
[0175] Prompt Sentence Examples
[0176] Example of conversation target information input: "Ryo, male, 25 years old, cheerful and sociable, university trip"
[0177] Example user message: "Ryo, how are you doing lately?"
[0178] Example response from the generative AI model: "It's been a while! Work has been busy lately..."
[0179] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0180] Program processing flow
[0181] Step 1: Collecting User Input
[0182] The user opens the chat app and accesses the "Conversation Target Information Input Screen." Next, they enter information about the person they are talking to, such as their name, gender, age, personality, and past memories. An example of the input information is "Ryo, male, 25 years old, cheerful and sociable, college trip." When the user clicks the send button, this information is sent to the device.
[0183] Input: User-entered information about the person you are speaking to
[0184] Output: Formatted information sent to the terminal
[0185] Step 2: Sending information from the device to the server
[0186] The device receives input information from the user and formats it into JSON format. This formatted data is securely sent to the server using the HTTPS protocol. For example, the data format sent might look like this: {"name": "Ryo", "gender": "Male", "age": "25", "personality": "cheerful and sociable", "memory": "university trip"}
[0187] Input: User-entered information
[0188] Output: JSON format data sent to the server
[0189] Step 3: Save information on the server and check its integrity
[0190] The server receives the information from the device and stores it in a database. Before saving, it checks whether the name, age, gender, personality, and past memories are entered correctly. For example, it checks whether the name is not empty, whether the age is a number, etc. After the check is complete, it stores the information in the database.
[0191] Input: JSON format data received from the terminal
[0192] Output: Information stored in the database
[0193] Step 4: Training the AI model
[0194] The server reads the subject's information from the database and provides it as input to a generative AI model (e.g., GPT-4). The generative AI model learns dialogue patterns that reflect the subject's personality and past memories. For example, it learns based on the information needed to imitate the behavior of a 25-year-old man named "Ryo" who has a "cheerful and sociable" personality.
[0195] Input: Subject information read from the database
[0196] Output: A trained generative AI model
[0197] Step 5: Start a chat session
[0198] To start a conversation with a conversation target "Ryo" in a chat app, the user selects "Ryo" from the target selection screen and clicks the Start Chat button. The device sends a chat session start request to the server. This request includes the user ID and the selected target ID. The server generates a new session ID, launches the trained generative AI model, and associates it with this session ID. This makes it possible to track interactions between a specific user and target.
[0199] Input: User-selected target information, chat session start request
[0200] Output: Session ID and launch of trained generative AI model
[0201] Step 6: Process messages from users
[0202] The user types a message on the chat screen. For example, "How are you doing lately?" The device reads the message and sends it to the server. This message is also formatted in JSON: {"sessionId": "12345", "message": "How are you doing lately?"}
[0203] Input: The message entered by the user
[0204] Output: The formatted message sent to the server
[0205] Step 7: Response Generation and Sending
[0206] The server receives messages from users and inputs them into the generative AI model. The generative AI model analyzes the messages and generates appropriate responses. For example, in response to a user's message, "How are you doing lately?", the model generates a response such as, "It's been a while! I've been busy at work lately..." The generated response is then sent back to the device. The device receives the response from the server and displays it on the chat screen. This allows the user to enter the next message.
[0207] Input: Message received from user, trained generative AI model
[0208] Output: The generated response and its display in the chat screen
[0209] (Application example 1)
[0210] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0211] In modern society, many people feel lonely and isolated, and there is a need to address this. Furthermore, many consumers feel that the shopping experience in physical stores is impersonal and inefficient. This can lead to stressful shopping activities. This invention aims to simultaneously alleviate this sense of loneliness and improve the shopping experience.
[0212] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0213] In this invention, the server includes a means for inputting information about a conversation target, a means for transmitting the information to the server, a means for the server to store the information in a database, a means for reading the target information from the database and training an AI model, a means for a user to have a dialogue with the AI model to support a shopping experience in a physical store, and a means for displaying the content of the dialogue to the user, thereby enabling the user to enjoy an intimate shopping experience through a conversation with a virtual store clerk.
[0214] "Conversational target" means information about a person with whom a user can have a virtual conversation.
[0215] "Server" refers to the computer system that stores information received from users and trains the AI model and generates dialogue.
[0216] "Database" refers to a database management system for storing and managing information on conversation partners.
[0217] "AI model" refers to an artificial intelligence model that has been trained to converse with users based on the personality and past memories of the person being spoken to.
[0218] "Learning" refers to the process by which an AI model generates conversation patterns based on information about the person being spoken to.
[0219] A "physical store" refers to a store that physically exists and where users visit to purchase products.
[0220] "Shopping experience" refers to the entire purchasing activity a user engages in at a physical store, including product selection, purchase, and in-store activities.
[0221] "Dialogue content" refers to the messages exchanged between the user and the AI model in real time.
[0222] "User" refers to an individual who uses the system to engage in virtual conversations.
[0223] The invention is a system that improves the shopping experience for users in physical stores through virtual conversations.
[0224] System configuration
[0225] First, the user accesses the system using a smartphone or head-mounted display. The user puts on the application and inputs information about the person they are talking to, such as the person's name, gender, age, personality, and past memories. This generates an individual AI model, ready to engage in real-time dialogue with the user.
[0226] Hardware and software used
[0227] Hardware:
[0228] Smartphone (e.g. Samsung Galaxy)
[0229] Head-mounted displays (e.g. Microsoft HoloLens)
[0230] software:
[0231] Natural Language Processing (NLP) libraries (e.g., spaCy)
[0232] HTTPS communication protocol
[0233] JSON Parser
[0234] Generative AI models (e.g., OpenAI GPT-4)
[0235] Database management system (e.g. MySQL)
[0236] Data processing and calculation
[0237] server:
[0238] 1. The server receives the conversation target information sent by the user. This information is sent in JSON format and securely stored in a database.
[0239] 2. The AI model is trained based on the information stored in the database, which generates responses that reflect the subject's personality and past memories.
[0240] Device:
[0241] 3. The device requests a chat session from the server when the user wants to start a conversation.
[0242] 4. Receive a response from the server and display it on the chat screen or head-mounted display.
[0243] user:
[0244] 5. While shopping in a physical store, users ask questions about products, which are sent to the server via their device.
[0245] 6. The server performs natural language analysis of the question and generates an appropriate response.
[0246] 7. The generated response is sent back to the terminal and displayed to the user.
[0247] Specific examples
[0248] Simulating a brick-and-mortar conversation:
[0249] User: "What do you think of this dress?"
[0250] Virtual Salesperson: "It's so pretty! I think it would be perfect for a party. Maybe something with a bit more frills would suit your style."
[0251] Example prompt sentence:
[0252] "Hi, I came in today to look at some new bags. Do you have any recommendations?"
[0253] "Which jacket do you think would suit me? What colors are popular these days?"
[0254] This system allows users to enjoy an intimate and personal shopping experience through conversations with virtual store associates. This invention is expected to dramatically improve the traditional shopping experience and increase user satisfaction.
[0255] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0256] Step 1:
[0257] The user launches the application using a smartphone or head-mounted display and inputs information about the person they are talking to, including their name, gender, age, personality, past memories, etc. This completes the necessary input data.
[0258] Step 2:
[0259] The terminal formats the conversation target information entered by the user into JSON format and sends it securely to the server using the HTTPS protocol. The input is information from the user, and the output is data to be sent to the server.
[0260] Step 3:
[0261] The server stores the received conversation target information in a database. When storing, the data integrity is checked to confirm the accuracy of the input data (e.g., name, age, gender). The input is the data sent from the device, and the output is the data stored in the database.
[0262] Step 4:
[0263] The server trains the generative AI model based on the conversation target's information stored in the database. Specifically, it extracts the target's personality, memories, and other characteristics and trains the AI model based on these. The input is the database information, and the output is the trained AI model.
[0264] Step 5:
[0265] While shopping in a physical store, the user starts a virtual conversation using a chat app. The user asks a question (prompt sentence) about the product to the virtual store clerk to assist with the purchase on the device where the content is displayed. The input is the question from the user.
[0266] Step 6:
[0267] The device sends the user's question to the server, which receives the question and analyzes it using a natural language processing library (e.g., spaCy). The input is the user's question, and the output is the analyzed question data.
[0268] Step 7:
[0269] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the user's question. For example, in response to the question "What do you think of this shirt?", it generates a response such as "It's very nice. It suits you." The input is the parsed question data, and the output is the generated response.
[0270] Step 8:
[0271] The server returns the generated response to the terminal, which displays the response on the chat screen or head-mounted display. The input is the response data from the server, and the output is the response message displayed to the user.
[0272] Step 9:
[0273] The user checks the displayed response and continues the conversation or asks the next question. This processing loop allows the user to continue the virtual conversation and improve the shopping experience. The input is the user's new question, and the output is the submission data for the new question.
[0274] Through these steps, users can have a more personal and fulfilling shopping experience through real-time interaction with a virtual store associate.
[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0276] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it enables more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition functions, and is implemented according to the following steps.
[0277] Overall overview
[0278] Users input information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. The application also provides a message input function to recognize the user's emotions, and an emotion engine that analyzes facial expressions and voice. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and emotion engine, preparing it for a virtual conversation with the user. When a user starts a chat, a message is sent from the device to the server, and the server uses the AI model and emotion engine to generate an appropriate response, which is then displayed on the user's device. In this way, the virtual conversation progresses in real time.
[0279] Program processing
[0280] User:
[0281] Enter information about the person you are talking to (such as name, gender, age, personality, past memories, etc.). For example, enter the name, age, gender, personality, etc. of your friend Ryo, and write in detail about memories you had with him (such as a college trip).
[0282] Check the information you entered and click the "Submit" button.
[0283] Device:
[0284] Format the conversation target information entered by the user into JSON format.
[0285] Information is sent securely to the server using the HTTPS protocol.
[0286] server:
[0287] Store the received information in the database. Check the integrity of the input data and check for any anomalies (e.g., name, age, gender, personality, memories, etc.).
[0288] The subject's information is read from the database, provided as input to the AI model, and the AI model is launched.
[0289] The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[0290] User:
[0291] Open the chat app and tap the "Start Conversation" button.
[0292] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[0293] Device:
[0294] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[0295] server:
[0296] Receives messages and emotional data from users. The message undergoes natural language analysis and the emotion engine recognizes the user's emotions.
[0297] Based on the recognized emotion and the results of natural language analysis, the AI model generates an appropriate response. For example, if the user's emotion is recognized as "sad," the tone of the response will be adjusted and a kind word such as "That must be tough. I'm rooting for you" will be added.
[0298] Sends the response to the user's terminal.
[0299] Device:
[0300] Receives a response from the server and displays it on the chat screen.
[0301] Specific examples
[0302] Example 1: Conversation with a friend
[0303] If a user wants to enjoy a virtual conversation with a college friend named "Ryo," they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, and then analyzes the user's emotions with an emotion engine, recreating natural, emotional conversations such as casual conversations and reminiscing. For example, if a user inputs "How are you doing lately?" and the emotion engine determines from the user's facial expression that they are "sad," the server uses the AI model to generate a response such as, "I've been busy at work lately, but I'm glad I got to talk to you. Did something happen?"
[0304] Example 2: Conversations with a deceased parent
[0305] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message saying, "Mom, do you remember me?" and the emotion engine recognizes the emotion "nostalgia," the server will generate a response such as, "Of course! I remember you. We used to cook together a lot when you were little."
[0306] In this way, the present invention provides a virtual interactive experience that allows users to reflect on past relationships and find healing and comfort, and can further enrich the experience by recognizing the user's emotions and generating more thoughtful responses.
[0307] The processing flow will be explained below.
[0308] Understood. Below, we will explain in detail the processing steps of the system that combines the emotion engine.
[0309] Step 1:
[0310] User:
[0311] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[0312] Check the information you entered and click the "Submit" button.
[0313] Step 2:
[0314] Device:
[0315] Format the conversation target information entered by the user into JSON format.
[0316] Information is sent securely to the server using the HTTPS protocol.
[0317] Step 3:
[0318] server:
[0319] The received information is stored in a database.
[0320] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[0321] Step 4:
[0322] server:
[0323] Read the subject's information from the database.
[0324] The read information is provided as input to the AI model, and the AI model is started.
[0325] Step 5:
[0326] server:
[0327] The AI model learns the subject's personality and past memories.
[0328] Generate patterns for interaction.
[0329] Step 6:
[0330] User:
[0331] Open the chat app and tap the "Start Conversation" button.
[0332] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[0333] Step 7:
[0334] Device:
[0335] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[0336] Step 8:
[0337] server:
[0338] Receive messages and sentiment data from users.
[0339] Messages are analyzed using natural language and the emotion engine recognizes the user's emotions.
[0340] Step 9:
[0341] server:
[0342] Based on the recognized emotions and the results of natural language analysis, the AI model generates an appropriate response.
[0343] For example, if the user's emotion is recognized as "sad," the tone of the response can be adjusted to include a kind comment such as, "That must be tough. I'm rooting for you."
[0344] Step 10:
[0345] server:
[0346] The generated response is sent to the user's terminal.
[0347] Step 11:
[0348] Device:
[0349] Receives a response from the server and displays it on the chat screen.
[0350] Step 12:
[0351] User:
[0352] Continuously typing messages (e.g., "Have you traveled anywhere recently?").
[0353] Step 13:
[0354] Device:
[0355] Continue sending each user message to the server.
[0356] Step 14:
[0357] server:
[0358] Each message is received, natural language analysis is performed, and the emotion engine analyzes the user's emotions.
[0359] Based on the user's sentiment and the content of the message, the AI model generates an appropriate response (e.g., "I went to Hokkaido last week. It was so much fun!").
[0360] Step 15:
[0361] Device:
[0362] Receives each response from the server and displays it to the user.
[0363] Step 16:
[0364] User:
[0365] Tap the "End Session" button to end the chat session.
[0366] Step 17:
[0367] Device:
[0368] Sends a request to the server to end the chat session.
[0369] Step 18:
[0370] server:
[0371] Ends the current chat session based on the session ID.
[0372] Archive session data to a database or delete it as needed.
[0373] This allows users to look back on past relationships and find emotional healing by utilizing an emotion engine and AI model.
[0374] Example 2
[0375] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0376] Conventional dialogue systems have the problem that when users virtually converse with estranged friends or deceased family members, the conversations are not emotionally rich and natural. Furthermore, they lack the ability to properly recognize the user's emotions and adjust responses accordingly. This makes it difficult for users to experience emotional healing and comfort.
[0377] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input information about a conversation target (such as name, gender, age, personality, and past memories); a means for formatting the information into JSON format and transmitting it to the server via HTTPS; a means for the server to store the information in a database; a means for reading the target information from the database and training a generative AI model and an emotion engine; a means for initiating a chat session and having a dialogue between the user and the AI model; an emotion recognition means for analyzing the user's facial expressions and voice and acquiring emotion data; a means for the server to generate an appropriate response using the AI model based on the content of the dialogue and the emotion recognition results; and a means for transmitting the response to the user's terminal and displaying it on the user's chat screen. This provides a virtual dialogue experience that allows the user to reflect on past relationships and find emotional comfort and security. Recognizing emotions and adjusting responses enables a richer dialogue experience.
[0378] "User" refers to an end user who uses the dialogue system.
[0379] A "conversational target" refers to a person (friend, family, etc.) with whom the user wishes to converse.
[0380] "Information" refers to data about the person you are speaking with, such as their name, gender, age, personality, and past memories.
[0381] "JSON format" refers to a data structure in JavaScript Object Notation, a widely used format for exchanging and storing data.
[0382] The "HTTPS protocol" stands for HyperText Transfer Protocol Secure, a protocol for ensuring secure communication, and encrypts data before sending and receiving it.
[0383] "Server" refers to a computer system that processes, stores, and analyzes information submitted by users.
[0384] A "database" refers to a system that systematically organizes and stores information.
[0385] A "generative AI model" refers to a dialogue response model generated based on a machine learning algorithm, designed to enable natural dialogue with users.
[0386] An "emotion engine" refers to a software system that analyzes a user's facial expressions and voice to recognize their emotions.
[0387] A "chat session" refers to a series of conversational interactions between a user and an AI model.
[0388] "Emotion recognition means" refers to a system that provides the functionality to analyze a user's facial expressions and voice and identify their emotions.
[0389] "Natural language analysis" refers to the processes and techniques aimed at interpreting and understanding the meaning of messages from users.
[0390] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it realizes more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition function, and is implemented according to the following steps.
[0391] Hardware and software used
[0392] Hardware: The mobile devices and computers used by users
[0393] Software: Chat application, generative AI model, emotion engine, database management system (DBMS)
[0394] Overall flow
[0395] The user inputs information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. Once input is complete, the device formats the information into JSON format and securely transmits it to the server using the HTTPS protocol. The server stores the received information in a database and verifies the integrity of the input data. The server then reads the person's information from the database and trains the generative AI model and emotion engine. This prepares the device for a virtual conversation.
[0396] When a user starts a chat session, a message is sent from the device to the server. At the same time, the emotion engine analyzes the user's facial expressions and voice to obtain emotional data. The server performs natural language analysis and emotion recognition based on the received message and emotional data. The generative AI model then generates an appropriate response, which is sent to the user's device. This process occurs in real time, allowing the user to experience a natural, emotionally rich virtual conversation.
[0397] Specific examples
[0398] Example 1: Conversation with a friend
[0399] When a user wants to enjoy a virtual conversation with a college friend they've lost touch with, they input detailed information about the friend (such as name, gender, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message asking "How are you doing lately?" and the emotion engine determines that the response is "sad," the server will generate a response such as "I've been busy lately, but I'm glad to be able to talk to you. Is something going on?" and send it to the user. This allows for realistic, emotionally rich dialogue.
[0400] Example 2: Conversations with a deceased parent
[0401] If a user wishes to have a virtual conversation with their deceased mother, they input detailed information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message such as "Mom, do you remember me?" and the emotion engine recognizes this as "nostalgic," the server will generate a response such as "Of course! I remember you. We used to cook together a lot when you were little," and send it to the user. This allows the user to relive past memories and find emotional healing.
[0402] Prompt Sentence Examples
[0403] "I'd like to talk to my friend {Name} about {Memories}. {Name} has {Personality Traits} and we last spoke {When did we last speak}."
[0404] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0405] Step 1:
[0406] User: Enter information about the person you are talking to
[0407] A user opens a chat application and enters details about the person they are talking to (such as name, gender, age, personality, and past memories). For example, they enter the name, gender (male), age (30), personality (cheerful and sociable), and past memories (such as photos they took together on a college trip) of their friend "Friend A." They confirm the information they entered and click the "Send" button.
[0408] Input: Information about the person you are speaking to
[0409] Output: Check input information and send operation
[0410] Step 2:
[0411] Terminal: Formatting and transmitting information
[0412] The device formats the conversation target information entered by the user into JSON format, specifically creating the following data structure:
[0413] json
[0414] {
[0415] "name": "Friend A",
[0416] "gender": "male",
[0417] "age": 30,
[0418] "personality": "cheerful and sociable",
[0419] "memories": "Photos we took together on a college trip"
[0420] }
[0421] It then sends the information securely to the server using the HTTPS protocol.
[0422] Input: Information about the person you are talking to entered by the user
[0423] Output: Sends data formatted in JSON format.
[0424] Step 3:
[0425] Server: storing and verifying information
[0426] The server stores the received information in a database. Before storing it, it checks the integrity of the information, for example, whether any required fields (name, gender, age, personality, memories) are missing. If the information is determined to be correct, it records it in the database as follows:
[0427] sql
[0428] INSERT INTO conversation_data (name, gender, age, personality, memories) VALUES ('Friend A', 'Male', 30, 'Cheerful and sociable', 'Photo taken together on a college trip');
[0429] Input: Received conversation target information
[0430] Output: Records in the database and confirmation results
[0431] Step 4:
[0432] Server: Preparing the AI model
[0433] The server reads the subject's information from the database and provides it as input to the generative AI model and emotion engine. Specifically, it lists information that has not yet been learned and feeds it to the model. The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[0434] Input: Subject information stored in the database
[0435] Output: Trained AI model and dialogue patterns
[0436] Step 5:
[0437] User:Start a chat session
[0438] The user opens the chat app and taps the "Start conversation" button. This is a button operation on the UI. They then type a message to their friend "Friend A" (e.g., "How are you doing lately?"). At the same time as they type the message, the emotion engine analyzes the user's facial expressions and voice in real time.
[0439] Input: Start chat session
[0440] Output: Messages and emotion data from users
[0441] Step 6:
[0442] Terminal: Sending messages and emotional data
[0443] The device sends the user's text message and the emotion data obtained from the emotion engine to the server in JSON format, generating the following data structure:
[0444] json
[0445] {
[0446] "message": "How are you doing lately?",
[0447] "emotion": "sad"
[0448] }
[0449] This data is transmitted securely using the HTTPS protocol.
[0450] Input: Message and emotion data from the user
[0451] Output: JSON data sent to the server
[0452] Step 7:
[0453] Server: Parses messages and generates responses
[0454] The server receives the user's message and emotional data and performs natural language analysis and emotion recognition. It uses natural language processing technology (e.g., NLTK or SpaCy) to analyze the message and recognize the emotional data. If the user's emotion is recognized as "sad," it adjusts the tone of the response. It uses a generative AI model to generate a response like this: "I've been busy lately, but I'm glad to talk to you. Is something going on?"
[0455] Input: Message and emotion data from the user
[0456] Output: The generated response message
[0457] Step 8:
[0458] Terminal: Display response
[0459] The device displays the response received from the server on the chat screen, and notifies the user in real time that a response has arrived.
[0460] Input: Response message from the server
[0461] Output: Display on the chat screen and notifications
[0462] (Application example 2)
[0463] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0464] In modern society, physical and emotional distance makes it difficult to reconnect with estranged friends or deceased family members. In particular, when users are experiencing emotional distress, there is a need for a way to find emotional healing through virtual conversations with these loved ones. However, conventional systems lack the means to recognize users' emotions in real time and generate appropriate responses, making it difficult to provide more natural and emotionally rich interactions.
[0465] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for saving information about the conversation target in a database, means for reading the target information from the database and training a generative AI model and an emotion analysis engine, and means for analyzing the user's facial expressions and voice and generating a response according to the emotion. This allows the user to have an emotionally rich and natural conversation through a virtual conversation with an estranged friend or a deceased family member, thereby providing emotional healing and a sense of security.
[0466] "Interviewee" refers to a person with whom you have a virtual conversation, such as a friend or family member.
[0467] A "generative AI model" is an artificial intelligence model that generates natural language based on input data.
[0468] An "emotion analysis engine" refers to software or algorithms that analyze and identify a user's emotions from their facial expressions and voice.
[0469] A "database" is a digital data storage system for storing and managing information about conversation partners and users.
[0470] A "server" is a computer system that processes information received from users and stores and manages it in a database.
[0471] "Chat session" refers to the interface through which a user and a generative AI model can converse, and the flow of that conversation.
[0472] "Natural language analysis" is the process of analyzing text entered by a user and understanding its meaning.
[0473] "Response" refers to the reply message that a generative AI model generates in response to user input.
[0474] "User" refers to the end user who uses this system to engage in conversations.
[0475] "Facial expression" refers to the element that allows a specific emotion to be read from the movement of the user's facial muscles.
[0476] "Audio" refers to sound data used to analyze emotions and content from a user's speech.
[0477] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion analysis engine, it achieves more emotionally rich dialogue. This system is implemented in the following steps.
[0478] 1. User input phase:
[0479] Users input information about the person they wish to have a virtual conversation with, including details such as name, gender, age, personality, past memories, etc. The interface for inputting this information is provided as an application on a smartphone or tablet.
[0480] 2. Data transmission phase:
[0481] The device converts the information entered by the user into JSON format and sends it to the server using the secure HTTPS protocol, which keeps the data secure.
[0482] 3. Server-side processing phase:
[0483] The server stores the received information in a database, reads the target person information from the database, and trains a generative AI model (e.g., OpenAI GPT-3) and an emotion analysis engine (e.g., Microsoft Azure Face API).
[0484] 4. Dialogue initiation phase:
[0485] When a user starts a chat session, the user's message and emotional data are sent to the server, which then performs natural language analysis and sentiment analysis. For example, a generative AI model analyzes the user's text data, and a sentiment analysis engine reads emotions from the user's facial expressions.
[0486] 5. Response generation phase:
[0487] The server uses a generative AI model based on the analysis results to generate an appropriate response, which takes into account the user's emotions. The generated response is sent to the device and displayed on the user's chat screen.
[0488] Hardware and software used
[0489] The main hardware used is smartphones and tablets. The software uses React Native as the front end, Node.js and Express as the back end, and MongoDB as the database. The sentiment analysis engine uses Microsoft Azure Face API, and the generative AI model uses OpenAI GPT-3.
[0490] Specific use cases
[0491] For example, suppose a user wants to enjoy a virtual conversation with a college friend named Ryo. The user enters Ryo's information (such as name, gender, age, personality, and memories) into the application. The server uses this information to train a generative AI model and an emotion analysis engine. If the user enters "How are you doing lately?" and the emotion analysis engine determines from the user's facial expression that they are "sad," the server will generate a response such as "I've been busy at work lately, but I'm glad I got to talk to you. Is something going on?"
[0492] Prompt Sentence Examples
[0493] I have a friend named Ryo. He's 25 years old and kind-hearted, but a bit mischievous. He remembers a trip he went on with his college club. Please respond to the following message as Ryo, expressing your feelings.
[0494] User: How are you doing lately?
[0495] Liao:
[0496] In this way, this invention provides a virtual conversational experience that allows users to reflect on past relationships and find healing and comfort. In addition, it can realize a richer experience by recognizing the user's emotions and generating more thoughtful responses.
[0497] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0498] Step 1:
[0499] The user uses a smartphone or tablet to enter information about the person they are talking to (such as name, gender, age, personality, past memories, etc.) on the application screen. The information entered by the user is converted into JSON format.
[0500] Input: User-entered information about the person you are speaking to
[0501] Output: Data converted to JSON format
[0502] Step 2:
[0503] The device securely transmits the JSON-formatted data to the server using the HTTPS protocol. Communication encryption prevents data from being intercepted by third parties.
[0504] Input: JSON format data
[0505] Output: Sending to server completed
[0506] Step 3:
[0507] The server analyzes the received JSON format data and saves it in the database. When saving, it checks the integrity of the data and eliminates invalid data.
[0508] Input: JSON format data sent to the server
[0509] Output: Data saved to database
[0510] Step 4:
[0511] The server reads the subject's information from the database, initializes the generative AI model (OpenAI GPT-3) and emotion analysis engine (Microsoft Azure Face API), and trains them.
[0512] Input: Subject information from the database
[0513] Output: Generative AI model and sentiment analysis engine initialized and trained
[0514] Step 5:
[0515] A user initiates a chat session in the application and types a message, while the device simultaneously captures the user's facial expressions with a camera and collects audio data.
[0516] Input: User-entered messages, facial expression data, and voice data
[0517] Output: Message and emotion data converted to JSON format
[0518] Step 6:
[0519] The collected messages and emotion data are then sent to the server using the HTTPS protocol, and the data is encrypted to protect the user's privacy.
[0520] Input: Message and emotion data in JSON format
[0521] Output: Sending to server completed
[0522] Step 7:
[0523] The server analyzes the received message and emotional data using natural language analysis and an emotion analysis engine. Specifically, the generative AI model analyzes the meaning of the message content, and the emotion analysis engine identifies the user's emotions from facial expressions and voice.
[0524] Input: Message and emotion data in JSON format sent to the server
[0525] Output: Analysis results
[0526] Step 8:
[0527] The server uses a generative AI model to generate an appropriate response based on the analysis results. For example, if the user's emotion is recognized as "sad," the server adjusts the tone of the response and generates words that will make the user feel lighter.
[0528] Input: Analysis results
[0529] Output: The generated response
[0530] Step 9:
[0531] The server sends the generated response to the terminal, which displays it on the chat screen. The user can then view the displayed response and enter their next message.
[0532] Input: Generated response
[0533] Output: Display on chat screen completed
[0534] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0535] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0536] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0537] [Second embodiment]
[0538] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0539] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0540] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0541] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0542] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0543] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0544] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0545] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0546] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0547] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0548] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0549] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0550] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system uses generative AI to simulate realistic conversations and is implemented according to the following steps:
[0551] Overall overview
[0552] First, a user enters information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, past memories, etc. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and prepares it for a virtual conversation with the user. When a user starts a chat, a message from the device is sent to the server, and the server uses the AI model to generate an appropriate response, which is then sent back to the user's device. In this way, the virtual conversation progresses in real time.
[0553] Program processing
[0554] User:
[0555] 1. Enter information about the person you are talking to. For example, enter the name, gender, age, etc. of your college friend Ryo, and write in detail about your memories with him (e.g., a college trip).
[0556] 2. Check the information you entered and click the send button.
[0557] Device:
[0558] 1. The entered target information is sent to the server. The data is formatted in JSON and sent securely via HTTPS.
[0559] server:
[0560] 1. Store the received information in a database. Check the data integrity, for example, to make sure that name, age, and gender are entered correctly.
[0561] 2. Based on the saved information, the AI model is activated. The target person information is read from the database and provided as input data to the AI model.
[0562] 3. Based on the stored information, the AI model learns the subject's personality and past memories.
[0563] User:
[0564] 1. Open the chat app and tap the button to start a conversation with the person you set (e.g., "Ryo").
[0565] Device:
[0566] 1. Send a chat session start request to the server. The request includes the user ID and target ID.
[0567] server:
[0568] 1. Start a chat session by generating a new session ID, launching the trained AI model, and associating it with the session ID.
[0569] 2. Receive a message from the user and have the AI model process it. For example, if a user types "How are you doing lately?", the AI model will analyze this message and generate an appropriate response.
[0570] 3. A response is generated and sent to the device. For example, a response such as "It's been a while! Work has been busy lately..." is generated.
[0571] Device:
[0572] 1. Receive a response from the server and display it on the chat screen.
[0573] Specific examples
[0574] Example 1: Conversation with a friend
[0575] When a user wants to have a virtual conversation with a university friend named Ryo, they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, which can then recreate natural conversations such as casual conversations and reminiscing. As the chat progresses, the server generates responses that are characteristic of Ryo in response to questions from the user, providing an experience that feels as if you are having a real conversation.
[0576] Example 2: Conversations with a deceased parent
[0577] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. When the user sends a message such as "Mom, do you remember me?", the server generates a response that reflects the mother's personality and past conversations, such as "Of course! I remember you. We often cooked together when you were little."
[0578] In this way, the present invention can provide a virtual interactive experience that allows users to reflect on past relationships and gain emotional healing and comfort.
[0579] The processing flow will be explained below.
[0580] Understood. The processing steps are explained in detail below.
[0581] Step 1:
[0582] User:
[0583] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[0584] Check the information you entered and click the "Submit" button.
[0585] Step 2:
[0586] Device:
[0587] Format the conversation target information entered by the user into JSON format.
[0588] Information is sent securely to the server using the HTTPS protocol.
[0589] Step 3:
[0590] server:
[0591] The received information is stored in a database.
[0592] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[0593] Step 4:
[0594] server:
[0595] Read the subject's information from the database.
[0596] The read information is provided as input to the AI model, and the AI model is started.
[0597] Step 5:
[0598] server:
[0599] The AI model learns the subject's personality and past memories.
[0600] Generate patterns for dialogue and prepare natural responses.
[0601] Step 6:
[0602] User:
[0603] Open the chat app and tap the "Start Conversation" button.
[0604] Step 7:
[0605] Device:
[0606] A chat session start request is sent to the server, including the user ID and target ID.
[0607] Step 8:
[0608] server:
[0609] Generates a new session ID and initializes the current chat session.
[0610] Associate the trained AI model with a session ID.
[0611] Step 9:
[0612] User:
[0613] Type your first message on the chat screen (e.g., "How are things going lately?").
[0614] Step 10:
[0615] Device:
[0616] Sends the user's message to the server. The message data is sent along with the session ID.
[0617] Step 11:
[0618] server:
[0619] Receive messages from users.
[0620] The message is subjected to natural language analysis, and an AI model generates an appropriate response based on the analysis results.
[0621] Step 12:
[0622] server:
[0623] The response generated by the AI model is sent to the user device (e.g., "It's been a while! Work has been busy lately...").
[0624] Step 13:
[0625] Device:
[0626] Receives a response from the server and displays it on the chat screen.
[0627] Step 14:
[0628] User:
[0629] Continue chatting by continually typing messages (e.g., "Have you traveled anywhere recently?").
[0630] Step 15:
[0631] Device:
[0632] It continues to send each of the user's messages to the server.
[0633] Step 16:
[0634] server:
[0635] Each message is received, natural language analysis is performed, and an appropriate response is generated by the AI model (e.g., "I went to Hokkaido last week. It was so much fun!").
[0636] Step 17:
[0637] Device:
[0638] Each response from the server is received and displayed to the user.
[0639] Step 18:
[0640] User:
[0641] Tap the "End Session" button to end the chat session.
[0642] Step 19:
[0643] Device:
[0644] Sends a request to the server to end the chat session.
[0645] Step 20:
[0646] server:
[0647] Ends the current chat session based on the session ID.
[0648] Archive session data to a database or delete it as needed.
[0649] This allows users to look back on past relationships and find emotional healing through virtual conversations.
[0650] Example 1
[0651] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0652] Loneliness and isolation are serious problems in modern society. In particular, when relationships with close friends and family become distant, or when communication with loved ones who have already passed away is lost, there are limited ways to fill the void. In such situations, users need sufficient support to reflect on past relationships and find emotional healing and security. However, current systems have difficulty providing a more realistic and emotionally charged conversational experience.
[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0654] In this invention, the server includes a means for storing information about the conversation target in a database, a means for reading the conversation target information from the database and training the generative AI model, and a means for starting a new chat session in which the user and the generative AI model converse. This makes it possible to generate conversation patterns that reflect the conversation target's personality and past memories, providing the user with a realistic and emotional conversation experience.
[0655] "Information about the person you are talking to" refers to detailed information about the person you are talking to, such as their name, gender, age, personality, and past memories, entered by the user.
[0656] "Server" refers to the computer system that stores information received from users in a database, trains the generative AI model, and manages chat sessions.
[0657] "Database" refers to a storage system that stores conversation target information and user information for later reference and use.
[0658] A "generative AI model" refers to an artificial intelligence model that learns based on a specific algorithm and generates natural-looking dialogue between a user and a conversation target.
[0659] A "chat session" refers to a series of interactions in which a user initiates a dialogue with a conversation partner and continuously exchanges messages.
[0660] "Session ID" means a unique identifier generated to identify a particular chat session.
[0661] "Natural language analysis" refers to language processing technology for understanding messages from users and generating appropriate responses.
[0662] "Terminal" refers to the device (e.g., smartphone or PC) that a user uses to enter information and communicate with a server.
[0663] "Response" refers to the message generated by the generative AI model and sent to the user.
[0664] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system utilizes generative AI models to simulate realistic conversations and includes the following key elements:
[0665] Hardware and Software Configuration
[0666] Device:
[0667] A device used by a user to input information and communicate with a server. Specifically, this applies to smartphones and PCs.
[0668] server:
[0669] It is a computer system that processes information received from users and generates responses using generative AI models. Its main functions include linking with databases, managing generative AI models, and natural language analysis.
[0670] Database:
[0671] It is a storage system for saving conversation target information and user information for later reference and use.
[0672] Generative AI models:
[0673] It is an artificial intelligence model that learns based on a specific algorithm (e.g., GPT-4) and generates natural dialogue between the user and the person being spoken to.
[0674] System Operation
[0675] user:
[0676] First, a user accesses a chat application using a terminal. Next, they input information about the person they are talking to. This information includes the person's name, gender, age, personality, and past memories. For example, a user can input memories of a friend named "Ryo," a 25-year-old man with a "cheerful and sociable" personality.
[0677] Device:
[0678] Once the user enters the information and clicks the submit button, the device formats this information into JSON format and sends it securely to the server using the HTTPS protocol.
[0679] server:
[0680] The server receives the entered information and stores it in a database. When saving, it checks the integrity of the information, such as name, age, and gender, to ensure it is entered correctly. Once the check is complete, the information is saved in the database.
[0681] The server then reads the target person's information from the database and provides it as input to the generative AI model, which then learns from their personality and past memories to prepare for a conversation with the user.
[0682] When a user starts a chat, the server generates a new session ID, activates the trained generative AI model, and associates it with the session ID. When a message from the user is sent from the device to the server, the server uses the generative AI model to generate an appropriate response and sends it back to the device. The generated response is displayed on the device's chat screen.
[0683] Specific examples
[0684] Scenario of a conversation with a friend
[0685] To enjoy a virtual conversation with a friend from college named "Ryo," a user inputs information about Ryo. For example, the user might enter his name as "Lee," his gender as "male," his age as "25," and mention a "college trip" as a memory. The server uses this information to train a generative AI model. When the user sends a message in a chat app asking "How are you doing lately?", the server generates a response such as "It's been a while! I've been busy at work lately..." and sends it back to the user.
[0686] Scenarios for conversations with deceased parents
[0687] If a user wishes to have a virtual conversation with their deceased mother, they enter detailed information about her (e.g., name, gender, age, personality, memories). The server uses this information to train a generative AI model. When the user sends a message saying, "Mom, do you remember me?", the server generates a response saying, "Of course! I remember you. We used to cook together a lot when you were little," and sends it back to the user.
[0688] The system allows users to reflect on past relationships and find healing and comfort. Using generative AI models, the dialogue progresses in real time, providing users with a highly natural and emotionally rich conversational experience.
[0689] Prompt Sentence Examples
[0690] Example of conversation target information input: "Ryo, male, 25 years old, cheerful and sociable, university trip"
[0691] Example user message: "Ryo, how are you doing lately?"
[0692] Example response from the generative AI model: "It's been a while! Work has been busy lately..."
[0693] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0694] Program processing flow
[0695] Step 1: Collecting User Input
[0696] The user opens the chat app and accesses the "Conversation Target Information Input Screen." Next, they enter information about the person they are talking to, such as their name, gender, age, personality, and past memories. An example of the input information is "Ryo, male, 25 years old, cheerful and sociable, college trip." When the user clicks the send button, this information is sent to the device.
[0697] Input: User-entered information about the person you are speaking to
[0698] Output: Formatted information sent to the terminal
[0699] Step 2: Sending information from the device to the server
[0700] The device receives input information from the user and formats it into JSON format. This formatted data is securely sent to the server using the HTTPS protocol. For example, the data format sent might look like this: {"name": "Ryo", "gender": "Male", "age": "25", "personality": "cheerful and sociable", "memory": "university trip"}
[0701] Input: User-entered information
[0702] Output: JSON format data sent to the server
[0703] Step 3: Save information on the server and check its integrity
[0704] The server receives the information from the device and stores it in a database. Before saving, it checks whether the name, age, gender, personality, and past memories are entered correctly. For example, it checks whether the name is not empty, whether the age is a number, etc. After the check is complete, it stores the information in the database.
[0705] Input: JSON format data received from the terminal
[0706] Output: Information stored in the database
[0707] Step 4: Training the AI model
[0708] The server reads the subject's information from the database and provides it as input to a generative AI model (e.g., GPT-4). The generative AI model learns dialogue patterns that reflect the subject's personality and past memories. For example, it learns based on the information needed to imitate the behavior of a 25-year-old man named "Ryo" who has a "cheerful and sociable" personality.
[0709] Input: Subject information read from the database
[0710] Output: A trained generative AI model
[0711] Step 5: Start a chat session
[0712] To start a conversation with a conversation target "Ryo" in a chat app, the user selects "Ryo" from the target selection screen and clicks the Start Chat button. The device sends a chat session start request to the server. This request includes the user ID and the selected target ID. The server generates a new session ID, launches the trained generative AI model, and associates it with this session ID. This makes it possible to track interactions between a specific user and target.
[0713] Input: User-selected target information, chat session start request
[0714] Output: Session ID and launch of trained generative AI model
[0715] Step 6: Process messages from users
[0716] The user types a message on the chat screen. For example, "How are you doing lately?" The device reads the message and sends it to the server. This message is also formatted in JSON: {"sessionId": "12345", "message": "How are you doing lately?"}
[0717] Input: The message entered by the user
[0718] Output: The formatted message sent to the server
[0719] Step 7: Response Generation and Sending
[0720] The server receives messages from users and inputs them into the generative AI model. The generative AI model analyzes the messages and generates appropriate responses. For example, in response to a user's message, "How are you doing lately?", the model generates a response such as, "It's been a while! I've been busy at work lately..." The generated response is then sent back to the device. The device receives the response from the server and displays it on the chat screen. This allows the user to enter the next message.
[0721] Input: Message received from user, trained generative AI model
[0722] Output: The generated response and its display in the chat screen
[0723] (Application example 1)
[0724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0725] In modern society, many people feel lonely and isolated, and there is a need to address this. Furthermore, many consumers feel that the shopping experience in physical stores is impersonal and inefficient. This can lead to stressful shopping activities. This invention aims to simultaneously alleviate this sense of loneliness and improve the shopping experience.
[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0727] In this invention, the server includes a means for inputting information about a conversation target, a means for transmitting the information to the server, a means for the server to store the information in a database, a means for reading the target information from the database and training an AI model, a means for a user to have a dialogue with the AI model to support a shopping experience in a physical store, and a means for displaying the content of the dialogue to the user, thereby enabling the user to enjoy an intimate shopping experience through a conversation with a virtual store clerk.
[0728] "Conversational target" means information about a person with whom a user can have a virtual conversation.
[0729] "Server" refers to the computer system that stores information received from users and trains the AI model and generates dialogue.
[0730] "Database" refers to a database management system for storing and managing information on conversation partners.
[0731] "AI model" refers to an artificial intelligence model that has been trained to converse with users based on the personality and past memories of the person being spoken to.
[0732] "Learning" refers to the process by which an AI model generates conversation patterns based on information about the person being spoken to.
[0733] A "physical store" refers to a store that physically exists and where users visit to purchase products.
[0734] "Shopping experience" refers to the entire purchasing activity a user engages in at a physical store, including product selection, purchase, and in-store activities.
[0735] "Dialogue content" refers to the messages exchanged between the user and the AI model in real time.
[0736] "User" refers to an individual who uses the system to engage in virtual conversations.
[0737] The invention is a system that improves the shopping experience for users in physical stores through virtual conversations.
[0738] System configuration
[0739] First, the user accesses the system using a smartphone or head-mounted display. The user puts on the application and inputs information about the person they are talking to, such as the person's name, gender, age, personality, and past memories. This generates an individual AI model, ready to engage in real-time dialogue with the user.
[0740] Hardware and software used
[0741] Hardware:
[0742] Smartphone (e.g. Samsung Galaxy)
[0743] Head-mounted displays (e.g. Microsoft HoloLens)
[0744] software:
[0745] Natural Language Processing (NLP) libraries (e.g., spaCy)
[0746] HTTPS communication protocol
[0747] JSON Parser
[0748] Generative AI models (e.g., OpenAI GPT-4)
[0749] Database management system (e.g. MySQL)
[0750] Data processing and calculation
[0751] server:
[0752] 1. The server receives the conversation target information sent by the user. This information is sent in JSON format and securely stored in a database.
[0753] 2. The AI model is trained based on the information stored in the database, which generates responses that reflect the subject's personality and past memories.
[0754] Device:
[0755] 3. The device requests a chat session from the server when the user wants to start a conversation.
[0756] 4. Receive a response from the server and display it on the chat screen or head-mounted display.
[0757] user:
[0758] 5. While shopping in a physical store, users ask questions about products, which are sent to the server via their device.
[0759] 6. The server performs natural language analysis of the question and generates an appropriate response.
[0760] 7. The generated response is sent back to the terminal and displayed to the user.
[0761] Specific examples
[0762] Simulating a brick-and-mortar conversation:
[0763] User: "What do you think of this dress?"
[0764] Virtual Salesperson: "It's so pretty! I think it would be perfect for a party. Maybe something with a bit more frills would suit your style."
[0765] Example prompt sentence:
[0766] "Hi, I came in today to look at some new bags. Do you have any recommendations?"
[0767] "Which jacket do you think would suit me? What colors are popular these days?"
[0768] This system allows users to enjoy an intimate and personal shopping experience through conversations with virtual store associates. This invention is expected to dramatically improve the traditional shopping experience and increase user satisfaction.
[0769] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0770] Step 1:
[0771] The user launches the application using a smartphone or head-mounted display and inputs information about the person they are talking to, including their name, gender, age, personality, past memories, etc. This completes the necessary input data.
[0772] Step 2:
[0773] The terminal formats the conversation target information entered by the user into JSON format and sends it securely to the server using the HTTPS protocol. The input is information from the user, and the output is data to be sent to the server.
[0774] Step 3:
[0775] The server stores the received conversation target information in a database. When storing, the data integrity is checked to confirm the accuracy of the input data (e.g., name, age, gender). The input is the data sent from the device, and the output is the data stored in the database.
[0776] Step 4:
[0777] The server trains the generative AI model based on the conversation target's information stored in the database. Specifically, it extracts the target's personality, memories, and other characteristics and trains the AI model based on these. The input is the database information, and the output is the trained AI model.
[0778] Step 5:
[0779] While shopping in a physical store, the user starts a virtual conversation using a chat app. The user asks a question (prompt sentence) about the product to the virtual store clerk to assist with the purchase on the device where the content is displayed. The input is the question from the user.
[0780] Step 6:
[0781] The device sends the user's question to the server, which receives the question and analyzes it using a natural language processing library (e.g., spaCy). The input is the user's question, and the output is the analyzed question data.
[0782] Step 7:
[0783] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the user's question. For example, in response to the question "What do you think of this shirt?", it generates a response such as "It's very nice. It suits you." The input is the parsed question data, and the output is the generated response.
[0784] Step 8:
[0785] The server returns the generated response to the terminal, which displays the response on the chat screen or head-mounted display. The input is the response data from the server, and the output is the response message displayed to the user.
[0786] Step 9:
[0787] The user checks the displayed response and continues the conversation or asks the next question. This processing loop allows the user to continue the virtual conversation and improve the shopping experience. The input is the user's new question, and the output is the submission data for the new question.
[0788] Through these steps, users can have a more personal and fulfilling shopping experience through real-time interaction with a virtual store associate.
[0789] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0790] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it enables more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition functions, and is implemented according to the following steps.
[0791] Overall overview
[0792] Users input information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. The application also provides a message input function to recognize the user's emotions, and an emotion engine that analyzes facial expressions and voice. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and emotion engine, preparing it for a virtual conversation with the user. When a user starts a chat, a message is sent from the device to the server, and the server uses the AI model and emotion engine to generate an appropriate response, which is then displayed on the user's device. In this way, the virtual conversation progresses in real time.
[0793] Program processing
[0794] User:
[0795] Enter information about the person you are talking to (such as name, gender, age, personality, past memories, etc.). For example, enter the name, age, gender, personality, etc. of your friend Ryo, and write in detail about memories you had with him (such as a college trip).
[0796] Check the information you entered and click the "Submit" button.
[0797] Device:
[0798] Format the conversation target information entered by the user into JSON format.
[0799] Information is sent securely to the server using the HTTPS protocol.
[0800] server:
[0801] Store the received information in the database. Check the integrity of the input data and check for any anomalies (e.g., name, age, gender, personality, memories, etc.).
[0802] The subject's information is read from the database, provided as input to the AI model, and the AI model is launched.
[0803] The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[0804] User:
[0805] Open the chat app and tap the "Start Conversation" button.
[0806] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[0807] Device:
[0808] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[0809] server:
[0810] Receives messages and emotional data from users. The message undergoes natural language analysis and the emotion engine recognizes the user's emotions.
[0811] Based on the recognized emotion and the results of natural language analysis, the AI model generates an appropriate response. For example, if the user's emotion is recognized as "sad," the tone of the response will be adjusted and a kind word such as "That must be tough. I'm rooting for you" will be added.
[0812] Sends the response to the user's terminal.
[0813] Device:
[0814] Receives a response from the server and displays it on the chat screen.
[0815] Specific examples
[0816] Example 1: Conversation with a friend
[0817] If a user wants to enjoy a virtual conversation with a college friend named "Ryo," they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, and then analyzes the user's emotions with an emotion engine, recreating natural, emotional conversations such as casual conversations and reminiscing. For example, if a user inputs "How are you doing lately?" and the emotion engine determines from the user's facial expression that they are "sad," the server uses the AI model to generate a response such as, "I've been busy at work lately, but I'm glad I got to talk to you. Did something happen?"
[0818] Example 2: Conversations with a deceased parent
[0819] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message saying, "Mom, do you remember me?" and the emotion engine recognizes the emotion "nostalgia," the server will generate a response such as, "Of course! I remember you. We used to cook together a lot when you were little."
[0820] In this way, the present invention provides a virtual interactive experience that allows users to reflect on past relationships and find healing and comfort, and can further enrich the experience by recognizing the user's emotions and generating more thoughtful responses.
[0821] The processing flow will be explained below.
[0822] Understood. Below, we will explain in detail the processing steps of the system that combines the emotion engine.
[0823] Step 1:
[0824] User:
[0825] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[0826] Check the information you entered and click the "Submit" button.
[0827] Step 2:
[0828] Device:
[0829] Format the conversation target information entered by the user into JSON format.
[0830] Information is sent securely to the server using the HTTPS protocol.
[0831] Step 3:
[0832] server:
[0833] The received information is stored in a database.
[0834] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[0835] Step 4:
[0836] server:
[0837] Read the subject's information from the database.
[0838] The read information is provided as input to the AI model, and the AI model is started.
[0839] Step 5:
[0840] server:
[0841] The AI model learns the subject's personality and past memories.
[0842] Generate patterns for interaction.
[0843] Step 6:
[0844] User:
[0845] Open the chat app and tap the "Start Conversation" button.
[0846] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[0847] Step 7:
[0848] Device:
[0849] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[0850] Step 8:
[0851] server:
[0852] Receive messages and sentiment data from users.
[0853] Messages are analyzed using natural language and the emotion engine recognizes the user's emotions.
[0854] Step 9:
[0855] server:
[0856] Based on the recognized emotions and the results of natural language analysis, the AI model generates an appropriate response.
[0857] For example, if the user's emotion is recognized as "sad," the tone of the response can be adjusted to include a kind comment such as, "That must be tough. I'm rooting for you."
[0858] Step 10:
[0859] server:
[0860] The generated response is sent to the user's terminal.
[0861] Step 11:
[0862] Device:
[0863] Receives a response from the server and displays it on the chat screen.
[0864] Step 12:
[0865] User:
[0866] Continuously typing messages (e.g., "Have you traveled anywhere recently?").
[0867] Step 13:
[0868] Device:
[0869] Continue sending each user message to the server.
[0870] Step 14:
[0871] server:
[0872] Each message is received, natural language analysis is performed, and the emotion engine analyzes the user's emotions.
[0873] Based on the user's sentiment and the content of the message, the AI model generates an appropriate response (e.g., "I went to Hokkaido last week. It was so much fun!").
[0874] Step 15:
[0875] Device:
[0876] Receives each response from the server and displays it to the user.
[0877] Step 16:
[0878] User:
[0879] Tap the "End Session" button to end the chat session.
[0880] Step 17:
[0881] Device:
[0882] Sends a request to the server to end the chat session.
[0883] Step 18:
[0884] server:
[0885] Ends the current chat session based on the session ID.
[0886] Archive session data to a database or delete it as needed.
[0887] This allows users to look back on past relationships and find emotional healing by utilizing an emotion engine and AI model.
[0888] Example 2
[0889] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0890] Conventional dialogue systems have the problem that when users virtually converse with estranged friends or deceased family members, the conversations are not emotionally rich and natural. Furthermore, they lack the ability to properly recognize the user's emotions and adjust responses accordingly. This makes it difficult for users to experience emotional healing and comfort.
[0891] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input information about a conversation target (such as name, gender, age, personality, and past memories); a means for formatting the information into JSON format and transmitting it to the server via HTTPS; a means for the server to store the information in a database; a means for reading the target information from the database and training a generative AI model and an emotion engine; a means for initiating a chat session and having a dialogue between the user and the AI model; an emotion recognition means for analyzing the user's facial expressions and voice and acquiring emotion data; a means for the server to generate an appropriate response using the AI model based on the content of the dialogue and the emotion recognition results; and a means for transmitting the response to the user's terminal and displaying it on the user's chat screen. This provides a virtual dialogue experience that allows the user to reflect on past relationships and find emotional comfort and security. Recognizing emotions and adjusting responses enables a richer dialogue experience.
[0892] "User" refers to an end user who uses the dialogue system.
[0893] A "conversational target" refers to a person (friend, family, etc.) with whom the user wishes to converse.
[0894] "Information" refers to data about the person you are speaking with, such as their name, gender, age, personality, and past memories.
[0895] "JSON format" refers to a data structure in JavaScript Object Notation, a widely used format for exchanging and storing data.
[0896] The "HTTPS protocol" stands for HyperText Transfer Protocol Secure, a protocol for ensuring secure communication, and encrypts data before sending and receiving it.
[0897] "Server" refers to a computer system that processes, stores, and analyzes information submitted by users.
[0898] A "database" refers to a system that systematically organizes and stores information.
[0899] A "generative AI model" refers to a dialogue response model generated based on a machine learning algorithm, designed to enable natural dialogue with users.
[0900] An "emotion engine" refers to a software system that analyzes a user's facial expressions and voice to recognize their emotions.
[0901] A "chat session" refers to a series of conversational interactions between a user and an AI model.
[0902] "Emotion recognition means" refers to a system that provides the functionality to analyze a user's facial expressions and voice and identify their emotions.
[0903] "Natural language analysis" refers to the processes and techniques aimed at interpreting and understanding the meaning of messages from users.
[0904] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it realizes more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition function, and is implemented according to the following steps.
[0905] Hardware and software used
[0906] Hardware: The mobile devices and computers used by users
[0907] Software: Chat application, generative AI model, emotion engine, database management system (DBMS)
[0908] Overall flow
[0909] The user inputs information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. Once input is complete, the device formats the information into JSON format and securely transmits it to the server using the HTTPS protocol. The server stores the received information in a database and verifies the integrity of the input data. The server then reads the person's information from the database and trains the generative AI model and emotion engine. This prepares the device for a virtual conversation.
[0910] When a user starts a chat session, a message is sent from the device to the server. At the same time, the emotion engine analyzes the user's facial expressions and voice to obtain emotional data. The server performs natural language analysis and emotion recognition based on the received message and emotional data. The generative AI model then generates an appropriate response, which is sent to the user's device. This process occurs in real time, allowing the user to experience a natural, emotionally rich virtual conversation.
[0911] Specific examples
[0912] Example 1: Conversation with a friend
[0913] When a user wants to enjoy a virtual conversation with a college friend they've lost touch with, they input detailed information about the friend (such as name, gender, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message asking "How are you doing lately?" and the emotion engine determines that the response is "sad," the server will generate a response such as "I've been busy lately, but I'm glad to be able to talk to you. Is something going on?" and send it to the user. This allows for realistic, emotionally rich dialogue.
[0914] Example 2: Conversations with a deceased parent
[0915] If a user wishes to have a virtual conversation with their deceased mother, they input detailed information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message such as "Mom, do you remember me?" and the emotion engine recognizes this as "nostalgic," the server will generate a response such as "Of course! I remember you. We used to cook together a lot when you were little," and send it to the user. This allows the user to relive past memories and find emotional healing.
[0916] Prompt Sentence Examples
[0917] "I'd like to talk to my friend {Name} about {Memories}. {Name} has {Personality Traits} and we last spoke {When did we last speak}."
[0918] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0919] Step 1:
[0920] User: Enter information about the person you are talking to
[0921] A user opens a chat application and enters details about the person they are talking to (such as name, gender, age, personality, and past memories). For example, they enter the name, gender (male), age (30), personality (cheerful and sociable), and past memories (such as photos they took together on a college trip) of their friend "Friend A." They confirm the information they entered and click the "Send" button.
[0922] Input: Information about the person you are speaking to
[0923] Output: Check input information and send operation
[0924] Step 2:
[0925] Terminal: Formatting and transmitting information
[0926] The device formats the conversation target information entered by the user into JSON format, specifically creating the following data structure:
[0927] json
[0928] {
[0929] "name": "Friend A",
[0930] "gender": "male",
[0931] "age": 30,
[0932] "personality": "cheerful and sociable",
[0933] "memories": "Photos we took together on a college trip"
[0934] }
[0935] It then sends the information securely to the server using the HTTPS protocol.
[0936] Input: Information about the person you are talking to entered by the user
[0937] Output: Sends data formatted in JSON format.
[0938] Step 3:
[0939] Server: storing and verifying information
[0940] The server stores the received information in a database. Before storing it, it checks the integrity of the information, for example, whether any required fields (name, gender, age, personality, memories) are missing. If the information is determined to be correct, it records it in the database as follows:
[0941] sql
[0942] INSERT INTO conversation_data (name, gender, age, personality, memories) VALUES ('Friend A', 'Male', 30, 'Cheerful and sociable', 'Photo taken together on a college trip');
[0943] Input: Received conversation target information
[0944] Output: Records in the database and confirmation results
[0945] Step 4:
[0946] Server: Preparing the AI model
[0947] The server reads the subject's information from the database and provides it as input to the generative AI model and emotion engine. Specifically, it lists information that has not yet been learned and feeds it to the model. The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[0948] Input: Subject information stored in the database
[0949] Output: Trained AI model and dialogue patterns
[0950] Step 5:
[0951] User:Start a chat session
[0952] The user opens the chat app and taps the "Start conversation" button. This is a button operation on the UI. They then type a message to their friend "Friend A" (e.g., "How are you doing lately?"). At the same time as they type the message, the emotion engine analyzes the user's facial expressions and voice in real time.
[0953] Input: Start chat session
[0954] Output: Messages and emotion data from users
[0955] Step 6:
[0956] Terminal: Sending messages and emotional data
[0957] The device sends the user's text message and the emotion data obtained from the emotion engine to the server in JSON format, generating the following data structure:
[0958] json
[0959] {
[0960] "message": "How are you doing lately?",
[0961] "emotion": "sad"
[0962] }
[0963] This data is transmitted securely using the HTTPS protocol.
[0964] Input: Message and emotion data from the user
[0965] Output: JSON data sent to the server
[0966] Step 7:
[0967] Server: Parses messages and generates responses
[0968] The server receives the user's message and emotional data and performs natural language analysis and emotion recognition. It uses natural language processing technology (e.g., NLTK or SpaCy) to analyze the message and recognize the emotional data. If the user's emotion is recognized as "sad," it adjusts the tone of the response. It uses a generative AI model to generate a response like this: "I've been busy lately, but I'm glad to talk to you. Is something going on?"
[0969] Input: Message and emotion data from the user
[0970] Output: The generated response message
[0971] Step 8:
[0972] Terminal: Display response
[0973] The device displays the response received from the server on the chat screen, and notifies the user in real time that a response has arrived.
[0974] Input: Response message from the server
[0975] Output: Display on the chat screen and notifications
[0976] (Application example 2)
[0977] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0978] In modern society, physical and emotional distance makes it difficult to reconnect with estranged friends or deceased family members. In particular, when users are experiencing emotional distress, there is a need for a way to find emotional healing through virtual conversations with these loved ones. However, conventional systems lack the means to recognize users' emotions in real time and generate appropriate responses, making it difficult to provide more natural and emotionally rich interactions.
[0979] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for saving information about the conversation target in a database, means for reading the target information from the database and training a generative AI model and an emotion analysis engine, and means for analyzing the user's facial expressions and voice and generating a response according to the emotion. This allows the user to have an emotionally rich and natural conversation through a virtual conversation with an estranged friend or a deceased family member, thereby providing emotional healing and a sense of security.
[0980] "Interviewee" refers to a person with whom you have a virtual conversation, such as a friend or family member.
[0981] A "generative AI model" is an artificial intelligence model that generates natural language based on input data.
[0982] An "emotion analysis engine" refers to software or algorithms that analyze and identify a user's emotions from their facial expressions and voice.
[0983] A "database" is a digital data storage system for storing and managing information about conversation partners and users.
[0984] A "server" is a computer system that processes information received from users and stores and manages it in a database.
[0985] "Chat session" refers to the interface through which a user and a generative AI model can converse, and the flow of that conversation.
[0986] "Natural language analysis" is the process of analyzing text entered by a user and understanding its meaning.
[0987] "Response" refers to the reply message that a generative AI model generates in response to user input.
[0988] "User" refers to the end user who uses this system to engage in conversations.
[0989] "Facial expression" refers to the element that allows a specific emotion to be read from the movement of the user's facial muscles.
[0990] "Audio" refers to sound data used to analyze emotions and content from a user's speech.
[0991] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion analysis engine, it achieves more emotionally rich dialogue. This system is implemented in the following steps.
[0992] 1. User input phase:
[0993] Users input information about the person they wish to have a virtual conversation with, including details such as name, gender, age, personality, past memories, etc. The interface for inputting this information is provided as an application on a smartphone or tablet.
[0994] 2. Data transmission phase:
[0995] The device converts the information entered by the user into JSON format and sends it to the server using the secure HTTPS protocol, which keeps the data secure.
[0996] 3. Server-side processing phase:
[0997] The server stores the received information in a database, reads the target person information from the database, and trains a generative AI model (e.g., OpenAI GPT-3) and an emotion analysis engine (e.g., Microsoft Azure Face API).
[0998] 4. Dialogue initiation phase:
[0999] When a user starts a chat session, the user's message and emotional data are sent to the server, which then performs natural language analysis and sentiment analysis. For example, a generative AI model analyzes the user's text data, and a sentiment analysis engine reads emotions from the user's facial expressions.
[1000] 5. Response generation phase:
[1001] The server uses a generative AI model based on the analysis results to generate an appropriate response, which takes into account the user's emotions. The generated response is sent to the device and displayed on the user's chat screen.
[1002] Hardware and software used
[1003] The main hardware used is smartphones and tablets. The software uses React Native as the front end, Node.js and Express as the back end, and MongoDB as the database. The sentiment analysis engine uses Microsoft Azure Face API, and the generative AI model uses OpenAI GPT-3.
[1004] Specific use cases
[1005] For example, suppose a user wants to enjoy a virtual conversation with a college friend named Ryo. The user enters Ryo's information (such as name, gender, age, personality, and memories) into the application. The server uses this information to train a generative AI model and an emotion analysis engine. If the user enters "How are you doing lately?" and the emotion analysis engine determines from the user's facial expression that they are "sad," the server will generate a response such as "I've been busy at work lately, but I'm glad I got to talk to you. Is something going on?"
[1006] Prompt Sentence Examples
[1007] I have a friend named Ryo. He's 25 years old and kind-hearted, but a bit mischievous. He remembers a trip he went on with his college club. Please respond to the following message as Ryo, expressing your feelings.
[1008] User: How are you doing lately?
[1009] Liao:
[1010] In this way, this invention provides a virtual conversational experience that allows users to reflect on past relationships and find healing and comfort. In addition, it can realize a richer experience by recognizing the user's emotions and generating more thoughtful responses.
[1011] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1012] Step 1:
[1013] The user uses a smartphone or tablet to enter information about the person they are talking to (such as name, gender, age, personality, past memories, etc.) on the application screen. The information entered by the user is converted into JSON format.
[1014] Input: User-entered information about the person you are speaking to
[1015] Output: Data converted to JSON format
[1016] Step 2:
[1017] The device securely transmits the JSON-formatted data to the server using the HTTPS protocol. Communication encryption prevents data from being intercepted by third parties.
[1018] Input: JSON format data
[1019] Output: Sending to server completed
[1020] Step 3:
[1021] The server analyzes the received JSON format data and saves it in the database. When saving, it checks the integrity of the data and eliminates invalid data.
[1022] Input: JSON format data sent to the server
[1023] Output: Data saved to database
[1024] Step 4:
[1025] The server reads the subject's information from the database, initializes the generative AI model (OpenAI GPT-3) and emotion analysis engine (Microsoft Azure Face API), and trains them.
[1026] Input: Subject information from the database
[1027] Output: Generative AI model and sentiment analysis engine initialized and trained
[1028] Step 5:
[1029] A user initiates a chat session in the application and types a message, while the device simultaneously captures the user's facial expressions with a camera and collects audio data.
[1030] Input: User-entered messages, facial expression data, and voice data
[1031] Output: Message and emotion data converted to JSON format
[1032] Step 6:
[1033] The collected messages and emotion data are then sent to the server using the HTTPS protocol, and the data is encrypted to protect the user's privacy.
[1034] Input: Message and emotion data in JSON format
[1035] Output: Sending to server completed
[1036] Step 7:
[1037] The server analyzes the received message and emotional data using natural language analysis and an emotion analysis engine. Specifically, the generative AI model analyzes the meaning of the message content, and the emotion analysis engine identifies the user's emotions from facial expressions and voice.
[1038] Input: Message and emotion data in JSON format sent to the server
[1039] Output: Analysis results
[1040] Step 8:
[1041] The server uses a generative AI model to generate an appropriate response based on the analysis results. For example, if the user's emotion is recognized as "sad," the server adjusts the tone of the response and generates words that will make the user feel lighter.
[1042] Input: Analysis results
[1043] Output: The generated response
[1044] Step 9:
[1045] The server sends the generated response to the terminal, which displays it on the chat screen. The user can then view the displayed response and enter their next message.
[1046] Input: Generated response
[1047] Output: Display on chat screen completed
[1048] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1049] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1050] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1051] [Third embodiment]
[1052] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1053] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1054] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1055] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1056] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1057] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1058] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1059] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1060] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1061] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1062] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1063] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1064] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system uses generative AI to simulate realistic conversations and is implemented according to the following steps:
[1065] Overall overview
[1066] First, a user enters information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, past memories, etc. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and prepares it for a virtual conversation with the user. When a user starts a chat, a message from the device is sent to the server, and the server uses the AI model to generate an appropriate response, which is then sent back to the user's device. In this way, the virtual conversation progresses in real time.
[1067] Program processing
[1068] User:
[1069] 1. Enter information about the person you are talking to. For example, enter the name, gender, age, etc. of your college friend Ryo, and write in detail about your memories with him (e.g., a college trip).
[1070] 2. Check the information you entered and click the send button.
[1071] Device:
[1072] 1. The entered target information is sent to the server. The data is formatted in JSON and sent securely via HTTPS.
[1073] server:
[1074] 1. Store the received information in a database. Check the data integrity, for example, to make sure that name, age, and gender are entered correctly.
[1075] 2. Based on the saved information, the AI model is activated. The target person information is read from the database and provided as input data to the AI model.
[1076] 3. Based on the stored information, the AI model learns the subject's personality and past memories.
[1077] User:
[1078] 1. Open the chat app and tap the button to start a conversation with the person you set (e.g., "Ryo").
[1079] Device:
[1080] 1. Send a chat session start request to the server. The request includes the user ID and target ID.
[1081] server:
[1082] 1. Start a chat session by generating a new session ID, launching the trained AI model, and associating it with the session ID.
[1083] 2. Receive a message from the user and have the AI model process it. For example, if a user types "How are you doing lately?", the AI model will analyze this message and generate an appropriate response.
[1084] 3. A response is generated and sent to the device. For example, a response such as "It's been a while! Work has been busy lately..." is generated.
[1085] Device:
[1086] 1. Receive a response from the server and display it on the chat screen.
[1087] Specific examples
[1088] Example 1: Conversation with a friend
[1089] When a user wants to have a virtual conversation with a university friend named Ryo, they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, which can then recreate natural conversations such as casual conversations and reminiscing. As the chat progresses, the server generates responses that are characteristic of Ryo in response to questions from the user, providing an experience that feels as if you are having a real conversation.
[1090] Example 2: Conversations with a deceased parent
[1091] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. When the user sends a message such as "Mom, do you remember me?", the server generates a response that reflects the mother's personality and past conversations, such as "Of course! I remember you. We often cooked together when you were little."
[1092] In this way, the present invention can provide a virtual interactive experience that allows users to reflect on past relationships and gain emotional healing and comfort.
[1093] The processing flow will be explained below.
[1094] Understood. The processing steps are explained in detail below.
[1095] Step 1:
[1096] User:
[1097] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[1098] Check the information you entered and click the "Submit" button.
[1099] Step 2:
[1100] Device:
[1101] Format the conversation target information entered by the user into JSON format.
[1102] Information is sent securely to the server using the HTTPS protocol.
[1103] Step 3:
[1104] server:
[1105] The received information is stored in a database.
[1106] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[1107] Step 4:
[1108] server:
[1109] Read the subject's information from the database.
[1110] The read information is provided as input to the AI model, and the AI model is started.
[1111] Step 5:
[1112] server:
[1113] The AI model learns the subject's personality and past memories.
[1114] Generate patterns for dialogue and prepare natural responses.
[1115] Step 6:
[1116] User:
[1117] Open the chat app and tap the "Start Conversation" button.
[1118] Step 7:
[1119] Device:
[1120] A chat session start request is sent to the server, including the user ID and target ID.
[1121] Step 8:
[1122] server:
[1123] Generates a new session ID and initializes the current chat session.
[1124] Associate the trained AI model with a session ID.
[1125] Step 9:
[1126] User:
[1127] Type your first message on the chat screen (e.g., "How are things going lately?").
[1128] Step 10:
[1129] Device:
[1130] Sends the user's message to the server. The message data is sent along with the session ID.
[1131] Step 11:
[1132] server:
[1133] Receive messages from users.
[1134] The message is subjected to natural language analysis, and an AI model generates an appropriate response based on the analysis results.
[1135] Step 12:
[1136] server:
[1137] The response generated by the AI model is sent to the user device (e.g., "It's been a while! Work has been busy lately...").
[1138] Step 13:
[1139] Device:
[1140] Receives a response from the server and displays it on the chat screen.
[1141] Step 14:
[1142] User:
[1143] Continue chatting by continually typing messages (e.g., "Have you traveled anywhere recently?").
[1144] Step 15:
[1145] Device:
[1146] It continues to send each of the user's messages to the server.
[1147] Step 16:
[1148] server:
[1149] Each message is received, natural language analysis is performed, and an appropriate response is generated by the AI model (e.g., "I went to Hokkaido last week. It was so much fun!").
[1150] Step 17:
[1151] Device:
[1152] Each response from the server is received and displayed to the user.
[1153] Step 18:
[1154] User:
[1155] Tap the "End Session" button to end the chat session.
[1156] Step 19:
[1157] Device:
[1158] Sends a request to the server to end the chat session.
[1159] Step 20:
[1160] server:
[1161] Ends the current chat session based on the session ID.
[1162] Archive session data to a database or delete it as needed.
[1163] This allows users to look back on past relationships and find emotional healing through virtual conversations.
[1164] Example 1
[1165] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1166] Loneliness and isolation are serious problems in modern society. In particular, when relationships with close friends and family become distant, or when communication with loved ones who have already passed away is lost, there are limited ways to fill the void. In such situations, users need sufficient support to reflect on past relationships and find emotional healing and security. However, current systems have difficulty providing a more realistic and emotionally charged conversational experience.
[1167] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1168] In this invention, the server includes a means for storing information about the conversation target in a database, a means for reading the conversation target information from the database and training the generative AI model, and a means for starting a new chat session in which the user and the generative AI model converse. This makes it possible to generate conversation patterns that reflect the conversation target's personality and past memories, providing the user with a realistic and emotional conversation experience.
[1169] "Information about the person you are talking to" refers to detailed information about the person you are talking to, such as their name, gender, age, personality, and past memories, entered by the user.
[1170] "Server" refers to the computer system that stores information received from users in a database, trains the generative AI model, and manages chat sessions.
[1171] "Database" refers to a storage system that stores conversation target information and user information for later reference and use.
[1172] A "generative AI model" refers to an artificial intelligence model that learns based on a specific algorithm and generates natural-looking dialogue between a user and a conversation target.
[1173] A "chat session" refers to a series of interactions in which a user initiates a dialogue with a conversation partner and continuously exchanges messages.
[1174] "Session ID" means a unique identifier generated to identify a particular chat session.
[1175] "Natural language analysis" refers to language processing technology for understanding messages from users and generating appropriate responses.
[1176] "Terminal" refers to the device (e.g., smartphone or PC) that a user uses to enter information and communicate with a server.
[1177] "Response" refers to the message generated by the generative AI model and sent to the user.
[1178] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system utilizes generative AI models to simulate realistic conversations and includes the following key elements:
[1179] Hardware and Software Configuration
[1180] Device:
[1181] A device used by a user to input information and communicate with a server. Specifically, this applies to smartphones and PCs.
[1182] server:
[1183] It is a computer system that processes information received from users and generates responses using generative AI models. Its main functions include linking with databases, managing generative AI models, and natural language analysis.
[1184] Database:
[1185] It is a storage system for saving conversation target information and user information for later reference and use.
[1186] Generative AI models:
[1187] It is an artificial intelligence model that learns based on a specific algorithm (e.g., GPT-4) and generates natural dialogue between the user and the person being spoken to.
[1188] System Operation
[1189] user:
[1190] First, a user accesses a chat application using a terminal. Next, they input information about the person they are talking to. This information includes the person's name, gender, age, personality, and past memories. For example, a user can input memories of a friend named "Ryo," a 25-year-old man with a "cheerful and sociable" personality.
[1191] Device:
[1192] Once the user enters the information and clicks the submit button, the device formats this information into JSON format and sends it securely to the server using the HTTPS protocol.
[1193] server:
[1194] The server receives the entered information and stores it in a database. When saving, it checks the integrity of the information, such as name, age, and gender, to ensure it is entered correctly. Once the check is complete, the information is saved in the database.
[1195] The server then reads the target person's information from the database and provides it as input to the generative AI model, which then learns from their personality and past memories to prepare for a conversation with the user.
[1196] When a user starts a chat, the server generates a new session ID, activates the trained generative AI model, and associates it with the session ID. When a message from the user is sent from the device to the server, the server uses the generative AI model to generate an appropriate response and sends it back to the device. The generated response is displayed on the device's chat screen.
[1197] Specific examples
[1198] Scenario of a conversation with a friend
[1199] To enjoy a virtual conversation with a friend from college named "Ryo," a user inputs information about Ryo. For example, the user might enter his name as "Lee," his gender as "male," his age as "25," and mention a "college trip" as a memory. The server uses this information to train a generative AI model. When the user sends a message in a chat app asking "How are you doing lately?", the server generates a response such as "It's been a while! I've been busy at work lately..." and sends it back to the user.
[1200] Scenarios for conversations with deceased parents
[1201] If a user wishes to have a virtual conversation with their deceased mother, they enter detailed information about her (e.g., name, gender, age, personality, memories). The server uses this information to train a generative AI model. When the user sends a message saying, "Mom, do you remember me?", the server generates a response saying, "Of course! I remember you. We used to cook together a lot when you were little," and sends it back to the user.
[1202] The system allows users to reflect on past relationships and find healing and comfort. Using generative AI models, the dialogue progresses in real time, providing users with a highly natural and emotionally rich conversational experience.
[1203] Prompt Sentence Examples
[1204] Example of conversation target information input: "Ryo, male, 25 years old, cheerful and sociable, university trip"
[1205] Example user message: "Ryo, how are you doing lately?"
[1206] Example response from the generative AI model: "It's been a while! Work has been busy lately..."
[1207] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1208] Program processing flow
[1209] Step 1: Collecting User Input
[1210] The user opens the chat app and accesses the "Conversation Target Information Input Screen." Next, they enter information about the person they are talking to, such as their name, gender, age, personality, and past memories. An example of the input information is "Ryo, male, 25 years old, cheerful and sociable, college trip." When the user clicks the send button, this information is sent to the device.
[1211] Input: User-entered information about the person you are speaking to
[1212] Output: Formatted information sent to the terminal
[1213] Step 2: Sending information from the device to the server
[1214] The device receives input information from the user and formats it into JSON format. This formatted data is securely sent to the server using the HTTPS protocol. For example, the data format sent might look like this: {"name": "Ryo", "gender": "Male", "age": "25", "personality": "cheerful and sociable", "memory": "university trip"}
[1215] Input: User-entered information
[1216] Output: JSON format data sent to the server
[1217] Step 3: Save information on the server and check its integrity
[1218] The server receives the information from the device and stores it in a database. Before saving, it checks whether the name, age, gender, personality, and past memories are entered correctly. For example, it checks whether the name is not empty, whether the age is a number, etc. After the check is complete, it stores the information in the database.
[1219] Input: JSON format data received from the terminal
[1220] Output: Information stored in the database
[1221] Step 4: Training the AI model
[1222] The server reads the subject's information from the database and provides it as input to a generative AI model (e.g., GPT-4). The generative AI model learns dialogue patterns that reflect the subject's personality and past memories. For example, it learns based on the information needed to imitate the behavior of a 25-year-old man named "Ryo" who has a "cheerful and sociable" personality.
[1223] Input: Subject information read from the database
[1224] Output: A trained generative AI model
[1225] Step 5: Start a chat session
[1226] To start a conversation with a conversation target "Ryo" in a chat app, the user selects "Ryo" from the target selection screen and clicks the Start Chat button. The device sends a chat session start request to the server. This request includes the user ID and the selected target ID. The server generates a new session ID, launches the trained generative AI model, and associates it with this session ID. This makes it possible to track interactions between a specific user and target.
[1227] Input: User-selected target information, chat session start request
[1228] Output: Session ID and launch of trained generative AI model
[1229] Step 6: Process messages from users
[1230] The user types a message on the chat screen. For example, "How are you doing lately?" The device reads the message and sends it to the server. This message is also formatted in JSON: {"sessionId": "12345", "message": "How are you doing lately?"}
[1231] Input: The message entered by the user
[1232] Output: The formatted message sent to the server
[1233] Step 7: Response Generation and Sending
[1234] The server receives messages from users and inputs them into the generative AI model. The generative AI model analyzes the messages and generates appropriate responses. For example, in response to a user's message, "How are you doing lately?", the model generates a response such as, "It's been a while! I've been busy at work lately..." The generated response is then sent back to the device. The device receives the response from the server and displays it on the chat screen. This allows the user to enter the next message.
[1235] Input: Message received from user, trained generative AI model
[1236] Output: The generated response and its display in the chat screen
[1237] (Application example 1)
[1238] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1239] In modern society, many people feel lonely and isolated, and there is a need to address this. Furthermore, many consumers feel that the shopping experience in physical stores is impersonal and inefficient. This can lead to stressful shopping activities. This invention aims to simultaneously alleviate this sense of loneliness and improve the shopping experience.
[1240] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1241] In this invention, the server includes a means for inputting information about a conversation target, a means for transmitting the information to the server, a means for the server to store the information in a database, a means for reading the target information from the database and training an AI model, a means for a user to have a dialogue with the AI model to support a shopping experience in a physical store, and a means for displaying the content of the dialogue to the user, thereby enabling the user to enjoy an intimate shopping experience through a conversation with a virtual store clerk.
[1242] "Conversational target" means information about a person with whom a user can have a virtual conversation.
[1243] "Server" refers to the computer system that stores information received from users and trains the AI model and generates dialogue.
[1244] "Database" refers to a database management system for storing and managing information on conversation partners.
[1245] "AI model" refers to an artificial intelligence model that has been trained to converse with users based on the personality and past memories of the person being spoken to.
[1246] "Learning" refers to the process by which an AI model generates conversation patterns based on information about the person being spoken to.
[1247] A "physical store" refers to a store that physically exists and where users visit to purchase products.
[1248] "Shopping experience" refers to the entire purchasing activity a user engages in at a physical store, including product selection, purchase, and in-store activities.
[1249] "Dialogue content" refers to the messages exchanged between the user and the AI model in real time.
[1250] "User" refers to an individual who uses the system to engage in virtual conversations.
[1251] The invention is a system that improves the shopping experience for users in physical stores through virtual conversations.
[1252] System configuration
[1253] First, the user accesses the system using a smartphone or head-mounted display. The user puts on the application and inputs information about the person they are talking to, such as the person's name, gender, age, personality, and past memories. This generates an individual AI model, ready to engage in real-time dialogue with the user.
[1254] Hardware and software used
[1255] Hardware:
[1256] Smartphone (e.g. Samsung Galaxy)
[1257] Head-mounted displays (e.g. Microsoft HoloLens)
[1258] software:
[1259] Natural Language Processing (NLP) libraries (e.g., spaCy)
[1260] HTTPS communication protocol
[1261] JSON Parser
[1262] Generative AI models (e.g., OpenAI GPT-4)
[1263] Database management system (e.g. MySQL)
[1264] Data processing and calculation
[1265] server:
[1266] 1. The server receives the conversation target information sent by the user. This information is sent in JSON format and securely stored in a database.
[1267] 2. The AI model is trained based on the information stored in the database, which generates responses that reflect the subject's personality and past memories.
[1268] Device:
[1269] 3. The device requests a chat session from the server when the user wants to start a conversation.
[1270] 4. Receive a response from the server and display it on the chat screen or head-mounted display.
[1271] user:
[1272] 5. While shopping in a physical store, users ask questions about products, which are sent to the server via their device.
[1273] 6. The server performs natural language analysis of the question and generates an appropriate response.
[1274] 7. The generated response is sent back to the terminal and displayed to the user.
[1275] Specific examples
[1276] Simulating a brick-and-mortar conversation:
[1277] User: "What do you think of this dress?"
[1278] Virtual Salesperson: "It's so pretty! I think it would be perfect for a party. Maybe something with a bit more frills would suit your style."
[1279] Example prompt sentence:
[1280] "Hi, I came in today to look at some new bags. Do you have any recommendations?"
[1281] "Which jacket do you think would suit me? What colors are popular these days?"
[1282] This system allows users to enjoy an intimate and personal shopping experience through conversations with virtual store associates. This invention is expected to dramatically improve the traditional shopping experience and increase user satisfaction.
[1283] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1284] Step 1:
[1285] The user launches the application using a smartphone or head-mounted display and inputs information about the person they are talking to, including their name, gender, age, personality, past memories, etc. This completes the necessary input data.
[1286] Step 2:
[1287] The terminal formats the conversation target information entered by the user into JSON format and sends it securely to the server using the HTTPS protocol. The input is information from the user, and the output is data to be sent to the server.
[1288] Step 3:
[1289] The server stores the received conversation target information in a database. When storing, the data integrity is checked to confirm the accuracy of the input data (e.g., name, age, gender). The input is the data sent from the device, and the output is the data stored in the database.
[1290] Step 4:
[1291] The server trains the generative AI model based on the conversation target's information stored in the database. Specifically, it extracts the target's personality, memories, and other characteristics and trains the AI model based on these. The input is the database information, and the output is the trained AI model.
[1292] Step 5:
[1293] While shopping in a physical store, the user starts a virtual conversation using a chat app. The user asks a question (prompt sentence) about the product to the virtual store clerk to assist with the purchase on the device where the content is displayed. The input is the question from the user.
[1294] Step 6:
[1295] The device sends the user's question to the server, which receives the question and analyzes it using a natural language processing library (e.g., spaCy). The input is the user's question, and the output is the analyzed question data.
[1296] Step 7:
[1297] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the user's question. For example, in response to the question "What do you think of this shirt?", it generates a response such as "It's very nice. It suits you." The input is the parsed question data, and the output is the generated response.
[1298] Step 8:
[1299] The server returns the generated response to the terminal, which displays the response on the chat screen or head-mounted display. The input is the response data from the server, and the output is the response message displayed to the user.
[1300] Step 9:
[1301] The user checks the displayed response and continues the conversation or asks the next question. This processing loop allows the user to continue the virtual conversation and improve the shopping experience. The input is the user's new question, and the output is the submission data for the new question.
[1302] Through these steps, users can have a more personal and fulfilling shopping experience through real-time interaction with a virtual store associate.
[1303] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1304] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it enables more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition functions, and is implemented according to the following steps.
[1305] Overall overview
[1306] Users input information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. The application also provides a message input function to recognize the user's emotions, and an emotion engine that analyzes facial expressions and voice. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and emotion engine, preparing it for a virtual conversation with the user. When a user starts a chat, a message is sent from the device to the server, and the server uses the AI model and emotion engine to generate an appropriate response, which is then displayed on the user's device. In this way, the virtual conversation progresses in real time.
[1307] Program processing
[1308] User:
[1309] Enter information about the person you are talking to (such as name, gender, age, personality, past memories, etc.). For example, enter the name, age, gender, personality, etc. of your friend Ryo, and write in detail about memories you had with him (such as a college trip).
[1310] Check the information you entered and click the "Submit" button.
[1311] Device:
[1312] Format the conversation target information entered by the user into JSON format.
[1313] Information is sent securely to the server using the HTTPS protocol.
[1314] server:
[1315] Store the received information in the database. Check the integrity of the input data and check for any anomalies (e.g., name, age, gender, personality, memories, etc.).
[1316] The subject's information is read from the database, provided as input to the AI model, and the AI model is launched.
[1317] The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[1318] User:
[1319] Open the chat app and tap the "Start Conversation" button.
[1320] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[1321] Device:
[1322] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[1323] server:
[1324] Receives messages and emotional data from users. The message undergoes natural language analysis and the emotion engine recognizes the user's emotions.
[1325] Based on the recognized emotion and the results of natural language analysis, the AI model generates an appropriate response. For example, if the user's emotion is recognized as "sad," the tone of the response will be adjusted and a kind word such as "That must be tough. I'm rooting for you" will be added.
[1326] Sends the response to the user's terminal.
[1327] Device:
[1328] Receives a response from the server and displays it on the chat screen.
[1329] Specific examples
[1330] Example 1: Conversation with a friend
[1331] If a user wants to enjoy a virtual conversation with a college friend named "Ryo," they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, and then analyzes the user's emotions with an emotion engine, recreating natural, emotional conversations such as casual conversations and reminiscing. For example, if a user inputs "How are you doing lately?" and the emotion engine determines from the user's facial expression that they are "sad," the server uses the AI model to generate a response such as, "I've been busy at work lately, but I'm glad I got to talk to you. Did something happen?"
[1332] Example 2: Conversations with a deceased parent
[1333] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message saying, "Mom, do you remember me?" and the emotion engine recognizes the emotion "nostalgia," the server will generate a response such as, "Of course! I remember you. We used to cook together a lot when you were little."
[1334] In this way, the present invention provides a virtual interactive experience that allows users to reflect on past relationships and find healing and comfort, and can further enrich the experience by recognizing the user's emotions and generating more thoughtful responses.
[1335] The processing flow will be explained below.
[1336] Understood. Below, we will explain in detail the processing steps of the system that combines the emotion engine.
[1337] Step 1:
[1338] User:
[1339] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[1340] Check the information you entered and click the "Submit" button.
[1341] Step 2:
[1342] Device:
[1343] Format the conversation target information entered by the user into JSON format.
[1344] Information is sent securely to the server using the HTTPS protocol.
[1345] Step 3:
[1346] server:
[1347] The received information is stored in a database.
[1348] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[1349] Step 4:
[1350] server:
[1351] Read the subject's information from the database.
[1352] The read information is provided as input to the AI model, and the AI model is started.
[1353] Step 5:
[1354] server:
[1355] The AI model learns the subject's personality and past memories.
[1356] Generate patterns for interaction.
[1357] Step 6:
[1358] User:
[1359] Open the chat app and tap the "Start Conversation" button.
[1360] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[1361] Step 7:
[1362] Device:
[1363] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[1364] Step 8:
[1365] server:
[1366] Receive messages and sentiment data from users.
[1367] Messages are analyzed using natural language and the emotion engine recognizes the user's emotions.
[1368] Step 9:
[1369] server:
[1370] Based on the recognized emotions and the results of natural language analysis, the AI model generates an appropriate response.
[1371] For example, if the user's emotion is recognized as "sad," the tone of the response can be adjusted to include a kind comment such as, "That must be tough. I'm rooting for you."
[1372] Step 10:
[1373] server:
[1374] The generated response is sent to the user's terminal.
[1375] Step 11:
[1376] Device:
[1377] Receives a response from the server and displays it on the chat screen.
[1378] Step 12:
[1379] User:
[1380] Continuously typing messages (e.g., "Have you traveled anywhere recently?").
[1381] Step 13:
[1382] Device:
[1383] Continue sending each user message to the server.
[1384] Step 14:
[1385] server:
[1386] Each message is received, natural language analysis is performed, and the emotion engine analyzes the user's emotions.
[1387] Based on the user's sentiment and the content of the message, the AI model generates an appropriate response (e.g., "I went to Hokkaido last week. It was so much fun!").
[1388] Step 15:
[1389] Device:
[1390] Receives each response from the server and displays it to the user.
[1391] Step 16:
[1392] User:
[1393] Tap the "End Session" button to end the chat session.
[1394] Step 17:
[1395] Device:
[1396] Sends a request to the server to end the chat session.
[1397] Step 18:
[1398] server:
[1399] Ends the current chat session based on the session ID.
[1400] Archive session data to a database or delete it as needed.
[1401] This allows users to look back on past relationships and find emotional healing by utilizing an emotion engine and AI model.
[1402] Example 2
[1403] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1404] Conventional dialogue systems have the problem that when users virtually converse with estranged friends or deceased family members, the conversations are not emotionally rich and natural. Furthermore, they lack the ability to properly recognize the user's emotions and adjust responses accordingly. This makes it difficult for users to experience emotional healing and comfort.
[1405] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input information about a conversation target (such as name, gender, age, personality, and past memories); a means for formatting the information into JSON format and transmitting it to the server via HTTPS; a means for the server to store the information in a database; a means for reading the target information from the database and training a generative AI model and an emotion engine; a means for initiating a chat session and having a dialogue between the user and the AI model; an emotion recognition means for analyzing the user's facial expressions and voice and acquiring emotion data; a means for the server to generate an appropriate response using the AI model based on the content of the dialogue and the emotion recognition results; and a means for transmitting the response to the user's terminal and displaying it on the user's chat screen. This provides a virtual dialogue experience that allows the user to reflect on past relationships and find emotional comfort and security. Recognizing emotions and adjusting responses enables a richer dialogue experience.
[1406] "User" refers to an end user who uses the dialogue system.
[1407] A "conversational target" refers to a person (friend, family, etc.) with whom the user wishes to converse.
[1408] "Information" refers to data about the person you are speaking with, such as their name, gender, age, personality, and past memories.
[1409] "JSON format" refers to a data structure in JavaScript Object Notation, a widely used format for exchanging and storing data.
[1410] The "HTTPS protocol" stands for HyperText Transfer Protocol Secure, a protocol for ensuring secure communication, and encrypts data before sending and receiving it.
[1411] "Server" refers to a computer system that processes, stores, and analyzes information submitted by users.
[1412] A "database" refers to a system that systematically organizes and stores information.
[1413] A "generative AI model" refers to a dialogue response model generated based on a machine learning algorithm, designed to enable natural dialogue with users.
[1414] An "emotion engine" refers to a software system that analyzes a user's facial expressions and voice to recognize their emotions.
[1415] A "chat session" refers to a series of conversational interactions between a user and an AI model.
[1416] "Emotion recognition means" refers to a system that provides the functionality to analyze a user's facial expressions and voice and identify their emotions.
[1417] "Natural language analysis" refers to the processes and techniques aimed at interpreting and understanding the meaning of messages from users.
[1418] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it realizes more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition function, and is implemented according to the following steps.
[1419] Hardware and software used
[1420] Hardware: The mobile devices and computers used by users
[1421] Software: Chat application, generative AI model, emotion engine, database management system (DBMS)
[1422] Overall flow
[1423] The user inputs information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. Once input is complete, the device formats the information into JSON format and securely transmits it to the server using the HTTPS protocol. The server stores the received information in a database and verifies the integrity of the input data. The server then reads the person's information from the database and trains the generative AI model and emotion engine. This prepares the device for a virtual conversation.
[1424] When a user starts a chat session, a message is sent from the device to the server. At the same time, the emotion engine analyzes the user's facial expressions and voice to obtain emotional data. The server performs natural language analysis and emotion recognition based on the received message and emotional data. The generative AI model then generates an appropriate response, which is sent to the user's device. This process occurs in real time, allowing the user to experience a natural, emotionally rich virtual conversation.
[1425] Specific examples
[1426] Example 1: Conversation with a friend
[1427] When a user wants to enjoy a virtual conversation with a college friend they've lost touch with, they input detailed information about the friend (such as name, gender, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message asking "How are you doing lately?" and the emotion engine determines that the response is "sad," the server will generate a response such as "I've been busy lately, but I'm glad to be able to talk to you. Is something going on?" and send it to the user. This allows for realistic, emotionally rich dialogue.
[1428] Example 2: Conversations with a deceased parent
[1429] If a user wishes to have a virtual conversation with their deceased mother, they input detailed information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message such as "Mom, do you remember me?" and the emotion engine recognizes this as "nostalgic," the server will generate a response such as "Of course! I remember you. We used to cook together a lot when you were little," and send it to the user. This allows the user to relive past memories and find emotional healing.
[1430] Prompt Sentence Examples
[1431] "I'd like to talk to my friend {Name} about {Memories}. {Name} has {Personality Traits} and we last spoke {When did we last speak}."
[1432] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1433] Step 1:
[1434] User: Enter information about the person you are talking to
[1435] A user opens a chat application and enters details about the person they are talking to (such as name, gender, age, personality, and past memories). For example, they enter the name, gender (male), age (30), personality (cheerful and sociable), and past memories (such as photos they took together on a college trip) of their friend "Friend A." They confirm the information they entered and click the "Send" button.
[1436] Input: Information about the person you are speaking to
[1437] Output: Check input information and send operation
[1438] Step 2:
[1439] Terminal: Formatting and transmitting information
[1440] The device formats the conversation target information entered by the user into JSON format, specifically creating the following data structure:
[1441] json
[1442] {
[1443] "name": "Friend A",
[1444] "gender": "male",
[1445] "age": 30,
[1446] "personality": "cheerful and sociable",
[1447] "memories": "Photos we took together on a college trip"
[1448] }
[1449] It then sends the information securely to the server using the HTTPS protocol.
[1450] Input: Information about the person you are talking to entered by the user
[1451] Output: Sends data formatted in JSON format.
[1452] Step 3:
[1453] Server: storing and verifying information
[1454] The server stores the received information in a database. Before storing it, it checks the integrity of the information, for example, whether any required fields (name, gender, age, personality, memories) are missing. If the information is determined to be correct, it records it in the database as follows:
[1455] sql
[1456] INSERT INTO conversation_data (name, gender, age, personality, memories) VALUES ('Friend A', 'Male', 30, 'Cheerful and sociable', 'Photo taken together on a college trip');
[1457] Input: Received conversation target information
[1458] Output: Records in the database and confirmation results
[1459] Step 4:
[1460] Server: Preparing the AI model
[1461] The server reads the subject's information from the database and provides it as input to the generative AI model and emotion engine. Specifically, it lists information that has not yet been learned and feeds it to the model. The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[1462] Input: Subject information stored in the database
[1463] Output: Trained AI model and dialogue patterns
[1464] Step 5:
[1465] User:Start a chat session
[1466] The user opens the chat app and taps the "Start conversation" button. This is a button operation on the UI. They then type a message to their friend "Friend A" (e.g., "How are you doing lately?"). At the same time as they type the message, the emotion engine analyzes the user's facial expressions and voice in real time.
[1467] Input: Start chat session
[1468] Output: Messages and emotion data from users
[1469] Step 6:
[1470] Terminal: Sending messages and emotional data
[1471] The device sends the user's text message and the emotion data obtained from the emotion engine to the server in JSON format, generating the following data structure:
[1472] json
[1473] {
[1474] "message": "How are you doing lately?",
[1475] "emotion": "sad"
[1476] }
[1477] This data is transmitted securely using the HTTPS protocol.
[1478] Input: Message and emotion data from the user
[1479] Output: JSON data sent to the server
[1480] Step 7:
[1481] Server: Parses messages and generates responses
[1482] The server receives the user's message and emotional data and performs natural language analysis and emotion recognition. It uses natural language processing technology (e.g., NLTK or SpaCy) to analyze the message and recognize the emotional data. If the user's emotion is recognized as "sad," it adjusts the tone of the response. It uses a generative AI model to generate a response like this: "I've been busy lately, but I'm glad to talk to you. Is something going on?"
[1483] Input: Message and emotion data from the user
[1484] Output: The generated response message
[1485] Step 8:
[1486] Terminal: Display response
[1487] The device displays the response received from the server on the chat screen, and notifies the user in real time that a response has arrived.
[1488] Input: Response message from the server
[1489] Output: Display on the chat screen and notifications
[1490] (Application example 2)
[1491] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1492] In modern society, physical and emotional distance makes it difficult to reconnect with estranged friends or deceased family members. In particular, when users are experiencing emotional distress, there is a need for a way to find emotional healing through virtual conversations with these loved ones. However, conventional systems lack the means to recognize users' emotions in real time and generate appropriate responses, making it difficult to provide more natural and emotionally rich interactions.
[1493] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for saving information about the conversation target in a database, means for reading the target information from the database and training a generative AI model and an emotion analysis engine, and means for analyzing the user's facial expressions and voice and generating a response according to the emotion. This allows the user to have an emotionally rich and natural conversation through a virtual conversation with an estranged friend or a deceased family member, thereby providing emotional healing and a sense of security.
[1494] "Interviewee" refers to a person with whom you have a virtual conversation, such as a friend or family member.
[1495] A "generative AI model" is an artificial intelligence model that generates natural language based on input data.
[1496] An "emotion analysis engine" refers to software or algorithms that analyze and identify a user's emotions from their facial expressions and voice.
[1497] A "database" is a digital data storage system for storing and managing information about conversation partners and users.
[1498] A "server" is a computer system that processes information received from users and stores and manages it in a database.
[1499] "Chat session" refers to the interface through which a user and a generative AI model can converse, and the flow of that conversation.
[1500] "Natural language analysis" is the process of analyzing text entered by a user and understanding its meaning.
[1501] "Response" refers to the reply message that a generative AI model generates in response to user input.
[1502] "User" refers to the end user who uses this system to engage in conversations.
[1503] "Facial expression" refers to the element that allows a specific emotion to be read from the movement of the user's facial muscles.
[1504] "Audio" refers to sound data used to analyze emotions and content from a user's speech.
[1505] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion analysis engine, it achieves more emotionally rich dialogue. This system is implemented in the following steps.
[1506] 1. User input phase:
[1507] Users input information about the person they wish to have a virtual conversation with, including details such as name, gender, age, personality, past memories, etc. The interface for inputting this information is provided as an application on a smartphone or tablet.
[1508] 2. Data transmission phase:
[1509] The device converts the information entered by the user into JSON format and sends it to the server using the secure HTTPS protocol, which keeps the data secure.
[1510] 3. Server-side processing phase:
[1511] The server stores the received information in a database, reads the target person information from the database, and trains a generative AI model (e.g., OpenAI GPT-3) and an emotion analysis engine (e.g., Microsoft Azure Face API).
[1512] 4. Dialogue initiation phase:
[1513] When a user starts a chat session, the user's message and emotional data are sent to the server, which then performs natural language analysis and sentiment analysis. For example, a generative AI model analyzes the user's text data, and a sentiment analysis engine reads emotions from the user's facial expressions.
[1514] 5. Response generation phase:
[1515] The server uses a generative AI model based on the analysis results to generate an appropriate response, which takes into account the user's emotions. The generated response is sent to the device and displayed on the user's chat screen.
[1516] Hardware and software used
[1517] The main hardware used is smartphones and tablets. The software uses React Native as the front end, Node.js and Express as the back end, and MongoDB as the database. The sentiment analysis engine uses Microsoft Azure Face API, and the generative AI model uses OpenAI GPT-3.
[1518] Specific use cases
[1519] For example, suppose a user wants to enjoy a virtual conversation with a college friend named Ryo. The user enters Ryo's information (such as name, gender, age, personality, and memories) into the application. The server uses this information to train a generative AI model and an emotion analysis engine. If the user enters "How are you doing lately?" and the emotion analysis engine determines from the user's facial expression that they are "sad," the server will generate a response such as "I've been busy at work lately, but I'm glad I got to talk to you. Is something going on?"
[1520] Prompt Sentence Examples
[1521] I have a friend named Ryo. He's 25 years old and kind-hearted, but a bit mischievous. He remembers a trip he went on with his college club. Please respond to the following message as Ryo, expressing your feelings.
[1522] User: How are you doing lately?
[1523] Liao:
[1524] In this way, this invention provides a virtual conversational experience that allows users to reflect on past relationships and find healing and comfort. In addition, it can realize a richer experience by recognizing the user's emotions and generating more thoughtful responses.
[1525] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1526] Step 1:
[1527] The user uses a smartphone or tablet to enter information about the person they are talking to (such as name, gender, age, personality, past memories, etc.) on the application screen. The information entered by the user is converted into JSON format.
[1528] Input: User-entered information about the person you are speaking to
[1529] Output: Data converted to JSON format
[1530] Step 2:
[1531] The device securely transmits the JSON-formatted data to the server using the HTTPS protocol. Communication encryption prevents data from being intercepted by third parties.
[1532] Input: JSON format data
[1533] Output: Sending to server completed
[1534] Step 3:
[1535] The server analyzes the received JSON format data and saves it in the database. When saving, it checks the integrity of the data and eliminates invalid data.
[1536] Input: JSON format data sent to the server
[1537] Output: Data saved to database
[1538] Step 4:
[1539] The server reads the subject's information from the database, initializes the generative AI model (OpenAI GPT-3) and emotion analysis engine (Microsoft Azure Face API), and trains them.
[1540] Input: Subject information from the database
[1541] Output: Generative AI model and sentiment analysis engine initialized and trained
[1542] Step 5:
[1543] A user initiates a chat session in the application and types a message, while the device simultaneously captures the user's facial expressions with a camera and collects audio data.
[1544] Input: User-entered messages, facial expression data, and voice data
[1545] Output: Message and emotion data converted to JSON format
[1546] Step 6:
[1547] The collected messages and emotion data are then sent to the server using the HTTPS protocol, and the data is encrypted to protect the user's privacy.
[1548] Input: Message and emotion data in JSON format
[1549] Output: Sending to server completed
[1550] Step 7:
[1551] The server analyzes the received message and emotional data using natural language analysis and an emotion analysis engine. Specifically, the generative AI model analyzes the meaning of the message content, and the emotion analysis engine identifies the user's emotions from facial expressions and voice.
[1552] Input: Message and emotion data in JSON format sent to the server
[1553] Output: Analysis results
[1554] Step 8:
[1555] The server uses a generative AI model to generate an appropriate response based on the analysis results. For example, if the user's emotion is recognized as "sad," the server adjusts the tone of the response and generates words that will make the user feel lighter.
[1556] Input: Analysis results
[1557] Output: The generated response
[1558] Step 9:
[1559] The server sends the generated response to the terminal, which displays it on the chat screen. The user can then view the displayed response and enter their next message.
[1560] Input: Generated response
[1561] Output: Display on chat screen completed
[1562] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1563] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1564] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1565] [Fourth embodiment]
[1566] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1567] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1568] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1569] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1570] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1571] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1572] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1573] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1574] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1575] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1576] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1577] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1578] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1579] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system uses generative AI to simulate realistic conversations and is implemented according to the following steps:
[1580] Overall overview
[1581] First, a user enters information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, past memories, etc. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and prepares it for a virtual conversation with the user. When a user starts a chat, a message from the device is sent to the server, and the server uses the AI model to generate an appropriate response, which is then sent back to the user's device. In this way, the virtual conversation progresses in real time.
[1582] Program processing
[1583] User:
[1584] 1. Enter information about the person you are talking to. For example, enter the name, gender, age, etc. of your college friend Ryo, and write in detail about your memories with him (e.g., a college trip).
[1585] 2. Check the information you entered and click the send button.
[1586] Device:
[1587] 1. The entered target information is sent to the server. The data is formatted in JSON and sent securely via HTTPS.
[1588] server:
[1589] 1. Store the received information in a database. Check the data integrity, for example, to make sure that name, age, and gender are entered correctly.
[1590] 2. Based on the saved information, the AI model is activated. The target person information is read from the database and provided as input data to the AI model.
[1591] 3. Based on the stored information, the AI model learns the subject's personality and past memories.
[1592] User:
[1593] 1. Open the chat app and tap the button to start a conversation with the person you set (e.g., "Ryo").
[1594] Device:
[1595] 1. Send a chat session start request to the server. The request includes the user ID and target ID.
[1596] server:
[1597] 1. Start a chat session by generating a new session ID, launching the trained AI model, and associating it with the session ID.
[1598] 2. Receive a message from the user and have the AI model process it. For example, if a user types "How are you doing lately?", the AI model will analyze this message and generate an appropriate response.
[1599] 3. A response is generated and sent to the device. For example, a response such as "It's been a while! Work has been busy lately..." is generated.
[1600] Device:
[1601] 1. Receive a response from the server and display it on the chat screen.
[1602] Specific examples
[1603] Example 1: Conversation with a friend
[1604] When a user wants to have a virtual conversation with a university friend named Ryo, they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, which can then recreate natural conversations such as casual conversations and reminiscing. As the chat progresses, the server generates responses that are characteristic of Ryo in response to questions from the user, providing an experience that feels as if you are having a real conversation.
[1605] Example 2: Conversations with a deceased parent
[1606] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. When the user sends a message such as "Mom, do you remember me?", the server generates a response that reflects the mother's personality and past conversations, such as "Of course! I remember you. We often cooked together when you were little."
[1607] In this way, the present invention can provide a virtual interactive experience that allows users to reflect on past relationships and gain emotional healing and comfort.
[1608] The processing flow will be explained below.
[1609] Understood. The processing steps are explained in detail below.
[1610] Step 1:
[1611] User:
[1612] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[1613] Check the information you entered and click the "Submit" button.
[1614] Step 2:
[1615] Device:
[1616] Format the conversation target information entered by the user into JSON format.
[1617] Information is sent securely to the server using the HTTPS protocol.
[1618] Step 3:
[1619] server:
[1620] The received information is stored in a database.
[1621] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[1622] Step 4:
[1623] server:
[1624] Read the subject's information from the database.
[1625] The read information is provided as input to the AI model, and the AI model is started.
[1626] Step 5:
[1627] server:
[1628] The AI model learns the subject's personality and past memories.
[1629] Generate patterns for dialogue and prepare natural responses.
[1630] Step 6:
[1631] User:
[1632] Open the chat app and tap the "Start Conversation" button.
[1633] Step 7:
[1634] Device:
[1635] A chat session start request is sent to the server, including the user ID and target ID.
[1636] Step 8:
[1637] server:
[1638] Generates a new session ID and initializes the current chat session.
[1639] Associate the trained AI model with a session ID.
[1640] Step 9:
[1641] User:
[1642] Type your first message on the chat screen (e.g., "How are things going lately?").
[1643] Step 10:
[1644] Device:
[1645] Sends the user's message to the server. The message data is sent along with the session ID.
[1646] Step 11:
[1647] server:
[1648] Receive messages from users.
[1649] The message is subjected to natural language analysis, and an AI model generates an appropriate response based on the analysis results.
[1650] Step 12:
[1651] server:
[1652] The response generated by the AI model is sent to the user device (e.g., "It's been a while! Work has been busy lately...").
[1653] Step 13:
[1654] Device:
[1655] Receives a response from the server and displays it on the chat screen.
[1656] Step 14:
[1657] User:
[1658] Continue chatting by continually typing messages (e.g., "Have you traveled anywhere recently?").
[1659] Step 15:
[1660] Device:
[1661] It continues to send each of the user's messages to the server.
[1662] Step 16:
[1663] server:
[1664] Each message is received, natural language analysis is performed, and an appropriate response is generated by the AI model (e.g., "I went to Hokkaido last week. It was so much fun!").
[1665] Step 17:
[1666] Device:
[1667] Each response from the server is received and displayed to the user.
[1668] Step 18:
[1669] User:
[1670] Tap the "End Session" button to end the chat session.
[1671] Step 19:
[1672] Device:
[1673] Sends a request to the server to end the chat session.
[1674] Step 20:
[1675] server:
[1676] Ends the current chat session based on the session ID.
[1677] Archive session data to a database or delete it as needed.
[1678] This allows users to look back on past relationships and find emotional healing through virtual conversations.
[1679] Example 1
[1680] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1681] Loneliness and isolation are serious problems in modern society. In particular, when relationships with close friends and family become distant, or when communication with loved ones who have already passed away is lost, there are limited ways to fill the void. In such situations, users need sufficient support to reflect on past relationships and find emotional healing and security. However, current systems have difficulty providing a more realistic and emotionally charged conversational experience.
[1682] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1683] In this invention, the server includes a means for storing information about the conversation target in a database, a means for reading the conversation target information from the database and training the generative AI model, and a means for starting a new chat session in which the user and the generative AI model converse. This makes it possible to generate conversation patterns that reflect the conversation target's personality and past memories, providing the user with a realistic and emotional conversation experience.
[1684] "Information about the person you are talking to" refers to detailed information about the person you are talking to, such as their name, gender, age, personality, and past memories, entered by the user.
[1685] "Server" refers to the computer system that stores information received from users in a database, trains the generative AI model, and manages chat sessions.
[1686] "Database" refers to a storage system that stores conversation target information and user information for later reference and use.
[1687] A "generative AI model" refers to an artificial intelligence model that learns based on a specific algorithm and generates natural-looking dialogue between a user and a conversation target.
[1688] A "chat session" refers to a series of interactions in which a user initiates a dialogue with a conversation partner and continuously exchanges messages.
[1689] "Session ID" means a unique identifier generated to identify a particular chat session.
[1690] "Natural language analysis" refers to language processing technology for understanding messages from users and generating appropriate responses.
[1691] "Terminal" refers to the device (e.g., smartphone or PC) that a user uses to enter information and communicate with a server.
[1692] "Response" refers to the message generated by the generative AI model and sent to the user.
[1693] The present invention is a system that helps users alleviate feelings of loneliness and isolation through virtual conversations with estranged friends and loved ones who have passed away. The system utilizes generative AI models to simulate realistic conversations and includes the following key elements:
[1694] Hardware and Software Configuration
[1695] Device:
[1696] A device used by a user to input information and communicate with a server. Specifically, this applies to smartphones and PCs.
[1697] server:
[1698] It is a computer system that processes information received from users and generates responses using generative AI models. Its main functions include linking with databases, managing generative AI models, and natural language analysis.
[1699] Database:
[1700] It is a storage system for saving conversation target information and user information for later reference and use.
[1701] Generative AI models:
[1702] It is an artificial intelligence model that learns based on a specific algorithm (e.g., GPT-4) and generates natural dialogue between the user and the person being spoken to.
[1703] System Operation
[1704] user:
[1705] First, a user accesses a chat application using a terminal. Next, they input information about the person they are talking to. This information includes the person's name, gender, age, personality, and past memories. For example, a user can input memories of a friend named "Ryo," a 25-year-old man with a "cheerful and sociable" personality.
[1706] Device:
[1707] Once the user enters the information and clicks the submit button, the device formats this information into JSON format and sends it securely to the server using the HTTPS protocol.
[1708] server:
[1709] The server receives the entered information and stores it in a database. When saving, it checks the integrity of the information, such as name, age, and gender, to ensure it is entered correctly. Once the check is complete, the information is saved in the database.
[1710] The server then reads the target person's information from the database and provides it as input to the generative AI model, which then learns from their personality and past memories to prepare for a conversation with the user.
[1711] When a user starts a chat, the server generates a new session ID, activates the trained generative AI model, and associates it with the session ID. When a message from the user is sent from the device to the server, the server uses the generative AI model to generate an appropriate response and sends it back to the device. The generated response is displayed on the device's chat screen.
[1712] Specific examples
[1713] Scenario of a conversation with a friend
[1714] To enjoy a virtual conversation with a friend from college named "Ryo," a user inputs information about Ryo. For example, the user might enter his name as "Lee," his gender as "male," his age as "25," and mention a "college trip" as a memory. The server uses this information to train a generative AI model. When the user sends a message in a chat app asking "How are you doing lately?", the server generates a response such as "It's been a while! I've been busy at work lately..." and sends it back to the user.
[1715] Scenarios for conversations with deceased parents
[1716] If a user wishes to have a virtual conversation with their deceased mother, they enter detailed information about her (e.g., name, gender, age, personality, memories). The server uses this information to train a generative AI model. When the user sends a message saying, "Mom, do you remember me?", the server generates a response saying, "Of course! I remember you. We used to cook together a lot when you were little," and sends it back to the user.
[1717] The system allows users to reflect on past relationships and find healing and comfort. Using generative AI models, the dialogue progresses in real time, providing users with a highly natural and emotionally rich conversational experience.
[1718] Prompt Sentence Examples
[1719] Example of conversation target information input: "Ryo, male, 25 years old, cheerful and sociable, university trip"
[1720] Example user message: "Ryo, how are you doing lately?"
[1721] Example response from the generative AI model: "It's been a while! Work has been busy lately..."
[1722] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1723] Program processing flow
[1724] Step 1: Collecting User Input
[1725] The user opens the chat app and accesses the "Conversation Target Information Input Screen." Next, they enter information about the person they are talking to, such as their name, gender, age, personality, and past memories. An example of the input information is "Ryo, male, 25 years old, cheerful and sociable, college trip." When the user clicks the send button, this information is sent to the device.
[1726] Input: User-entered information about the person you are speaking to
[1727] Output: Formatted information sent to the terminal
[1728] Step 2: Sending information from the device to the server
[1729] The device receives input information from the user and formats it into JSON format. This formatted data is securely sent to the server using the HTTPS protocol. For example, the data format sent might look like this: {"name": "Ryo", "gender": "Male", "age": "25", "personality": "cheerful and sociable", "memory": "university trip"}
[1730] Input: User-entered information
[1731] Output: JSON format data sent to the server
[1732] Step 3: Save information on the server and check its integrity
[1733] The server receives the information from the device and stores it in a database. Before saving, it checks whether the name, age, gender, personality, and past memories are entered correctly. For example, it checks whether the name is not empty, whether the age is a number, etc. After the check is complete, it stores the information in the database.
[1734] Input: JSON format data received from the terminal
[1735] Output: Information stored in the database
[1736] Step 4: Training the AI model
[1737] The server reads the subject's information from the database and provides it as input to a generative AI model (e.g., GPT-4). The generative AI model learns dialogue patterns that reflect the subject's personality and past memories. For example, it learns based on the information needed to imitate the behavior of a 25-year-old man named "Ryo" who has a "cheerful and sociable" personality.
[1738] Input: Subject information read from the database
[1739] Output: A trained generative AI model
[1740] Step 5: Start a chat session
[1741] To start a conversation with a conversation target "Ryo" in a chat app, the user selects "Ryo" from the target selection screen and clicks the Start Chat button. The device sends a chat session start request to the server. This request includes the user ID and the selected target ID. The server generates a new session ID, launches the trained generative AI model, and associates it with this session ID. This makes it possible to track interactions between a specific user and target.
[1742] Input: User-selected target information, chat session start request
[1743] Output: Session ID and launch of trained generative AI model
[1744] Step 6: Process messages from users
[1745] The user types a message on the chat screen. For example, "How are you doing lately?" The device reads the message and sends it to the server. This message is also formatted in JSON: {"sessionId": "12345", "message": "How are you doing lately?"}
[1746] Input: The message entered by the user
[1747] Output: The formatted message sent to the server
[1748] Step 7: Response Generation and Sending
[1749] The server receives messages from users and inputs them into the generative AI model. The generative AI model analyzes the messages and generates appropriate responses. For example, in response to a user's message, "How are you doing lately?", the model generates a response such as, "It's been a while! I've been busy at work lately..." The generated response is then sent back to the device. The device receives the response from the server and displays it on the chat screen. This allows the user to enter the next message.
[1750] Input: Message received from user, trained generative AI model
[1751] Output: The generated response and its display in the chat screen
[1752] (Application example 1)
[1753] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1754] In modern society, many people feel lonely and isolated, and there is a need to address this. Furthermore, many consumers feel that the shopping experience in physical stores is impersonal and inefficient. This can lead to stressful shopping activities. This invention aims to simultaneously alleviate this sense of loneliness and improve the shopping experience.
[1755] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1756] In this invention, the server includes a means for inputting information about a conversation target, a means for transmitting the information to the server, a means for the server to store the information in a database, a means for reading the target information from the database and training an AI model, a means for a user to have a dialogue with the AI model to support a shopping experience in a physical store, and a means for displaying the content of the dialogue to the user, thereby enabling the user to enjoy an intimate shopping experience through a conversation with a virtual store clerk.
[1757] "Conversational target" means information about a person with whom a user can have a virtual conversation.
[1758] "Server" refers to the computer system that stores information received from users and trains the AI model and generates dialogue.
[1759] "Database" refers to a database management system for storing and managing information on conversation partners.
[1760] "AI model" refers to an artificial intelligence model that has been trained to converse with users based on the personality and past memories of the person being spoken to.
[1761] "Learning" refers to the process by which an AI model generates conversation patterns based on information about the person being spoken to.
[1762] A "physical store" refers to a store that physically exists and where users visit to purchase products.
[1763] "Shopping experience" refers to the entire purchasing activity a user engages in at a physical store, including product selection, purchase, and in-store activities.
[1764] "Dialogue content" refers to the messages exchanged between the user and the AI model in real time.
[1765] "User" refers to an individual who uses the system to engage in virtual conversations.
[1766] The invention is a system that improves the shopping experience for users in physical stores through virtual conversations.
[1767] System configuration
[1768] First, the user accesses the system using a smartphone or head-mounted display. The user puts on the application and inputs information about the person they are talking to, such as the person's name, gender, age, personality, and past memories. This generates an individual AI model, ready to engage in real-time dialogue with the user.
[1769] Hardware and software used
[1770] Hardware:
[1771] Smartphone (e.g. Samsung Galaxy)
[1772] Head-mounted displays (e.g. Microsoft HoloLens)
[1773] software:
[1774] Natural Language Processing (NLP) libraries (e.g., spaCy)
[1775] HTTPS communication protocol
[1776] JSON Parser
[1777] Generative AI models (e.g., OpenAI GPT-4)
[1778] Database management system (e.g. MySQL)
[1779] Data processing and calculation
[1780] server:
[1781] 1. The server receives the conversation target information sent by the user. This information is sent in JSON format and securely stored in a database.
[1782] 2. The AI model is trained based on the information stored in the database, which generates responses that reflect the subject's personality and past memories.
[1783] Device:
[1784] 3. The device requests a chat session from the server when the user wants to start a conversation.
[1785] 4. Receive a response from the server and display it on the chat screen or head-mounted display.
[1786] user:
[1787] 5. While shopping in a physical store, users ask questions about products, which are sent to the server via their device.
[1788] 6. The server performs natural language analysis of the question and generates an appropriate response.
[1789] 7. The generated response is sent back to the terminal and displayed to the user.
[1790] Specific examples
[1791] Simulating a brick-and-mortar conversation:
[1792] User: "What do you think of this dress?"
[1793] Virtual Salesperson: "It's so pretty! I think it would be perfect for a party. Maybe something with a bit more frills would suit your style."
[1794] Example prompt sentence:
[1795] "Hi, I came in today to look at some new bags. Do you have any recommendations?"
[1796] "Which jacket do you think would suit me? What colors are popular these days?"
[1797] This system allows users to enjoy an intimate and personal shopping experience through conversations with virtual store associates. This invention is expected to dramatically improve the traditional shopping experience and increase user satisfaction.
[1798] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1799] Step 1:
[1800] The user launches the application using a smartphone or head-mounted display and inputs information about the person they are talking to, including their name, gender, age, personality, past memories, etc. This completes the necessary input data.
[1801] Step 2:
[1802] The terminal formats the conversation target information entered by the user into JSON format and sends it securely to the server using the HTTPS protocol. The input is information from the user, and the output is data to be sent to the server.
[1803] Step 3:
[1804] The server stores the received conversation target information in a database. When storing, the data integrity is checked to confirm the accuracy of the input data (e.g., name, age, gender). The input is the data sent from the device, and the output is the data stored in the database.
[1805] Step 4:
[1806] The server trains the generative AI model based on the conversation target's information stored in the database. Specifically, it extracts the target's personality, memories, and other characteristics and trains the AI model based on these. The input is the database information, and the output is the trained AI model.
[1807] Step 5:
[1808] While shopping in a physical store, the user starts a virtual conversation using a chat app. The user asks a question (prompt sentence) about the product to the virtual store clerk to assist with the purchase on the device where the content is displayed. The input is the question from the user.
[1809] Step 6:
[1810] The device sends the user's question to the server, which receives the question and analyzes it using a natural language processing library (e.g., spaCy). The input is the user's question, and the output is the analyzed question data.
[1811] Step 7:
[1812] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the user's question. For example, in response to the question "What do you think of this shirt?", it generates a response such as "It's very nice. It suits you." The input is the parsed question data, and the output is the generated response.
[1813] Step 8:
[1814] The server returns the generated response to the terminal, which displays the response on the chat screen or head-mounted display. The input is the response data from the server, and the output is the response message displayed to the user.
[1815] Step 9:
[1816] The user checks the displayed response and continues the conversation or asks the next question. This processing loop allows the user to continue the virtual conversation and improve the shopping experience. The input is the user's new question, and the output is the submission data for the new question.
[1817] Through these steps, users can have a more personal and fulfilling shopping experience through real-time interaction with a virtual store associate.
[1818] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1819] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it enables more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition functions, and is implemented according to the following steps.
[1820] Overall overview
[1821] Users input information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. The application also provides a message input function to recognize the user's emotions, and an emotion engine that analyzes facial expressions and voice. The device then sends this information to a server, which stores it in a database. The server uses this information to train an AI model and emotion engine, preparing it for a virtual conversation with the user. When a user starts a chat, a message is sent from the device to the server, and the server uses the AI model and emotion engine to generate an appropriate response, which is then displayed on the user's device. In this way, the virtual conversation progresses in real time.
[1822] Program processing
[1823] User:
[1824] Enter information about the person you are talking to (such as name, gender, age, personality, past memories, etc.). For example, enter the name, age, gender, personality, etc. of your friend Ryo, and write in detail about memories you had with him (such as a college trip).
[1825] Check the information you entered and click the "Submit" button.
[1826] Device:
[1827] Format the conversation target information entered by the user into JSON format.
[1828] Information is sent securely to the server using the HTTPS protocol.
[1829] server:
[1830] Store the received information in the database. Check the integrity of the input data and check for any anomalies (e.g., name, age, gender, personality, memories, etc.).
[1831] The subject's information is read from the database, provided as input to the AI model, and the AI model is launched.
[1832] The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[1833] User:
[1834] Open the chat app and tap the "Start Conversation" button.
[1835] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[1836] Device:
[1837] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[1838] server:
[1839] Receives messages and emotional data from users. The message undergoes natural language analysis and the emotion engine recognizes the user's emotions.
[1840] Based on the recognized emotion and the results of natural language analysis, the AI model generates an appropriate response. For example, if the user's emotion is recognized as "sad," the tone of the response will be adjusted and a kind word such as "That must be tough. I'm rooting for you" will be added.
[1841] Sends the response to the user's terminal.
[1842] Device:
[1843] Receives a response from the server and displays it on the chat screen.
[1844] Specific examples
[1845] Example 1: Conversation with a friend
[1846] If a user wants to enjoy a virtual conversation with a college friend named "Ryo," they input information about Ryo (such as his name, gender, age, personality, and memories). The server uses this information to train an AI model, and then analyzes the user's emotions with an emotion engine, recreating natural, emotional conversations such as casual conversations and reminiscing. For example, if a user inputs "How are you doing lately?" and the emotion engine determines from the user's facial expression that they are "sad," the server uses the AI model to generate a response such as, "I've been busy at work lately, but I'm glad I got to talk to you. Did something happen?"
[1847] Example 2: Conversations with a deceased parent
[1848] If a user wants to have a virtual conversation with their deceased mother, they input information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message saying, "Mom, do you remember me?" and the emotion engine recognizes the emotion "nostalgia," the server will generate a response such as, "Of course! I remember you. We used to cook together a lot when you were little."
[1849] In this way, the present invention provides a virtual interactive experience that allows users to reflect on past relationships and find healing and comfort, and can further enrich the experience by recognizing the user's emotions and generating more thoughtful responses.
[1850] The processing flow will be explained below.
[1851] Understood. Below, we will explain in detail the processing steps of the system that combines the emotion engine.
[1852] Step 1:
[1853] User:
[1854] Enter information about the person you are talking to (name, gender, age, personality, past memories, etc.).
[1855] Check the information you entered and click the "Submit" button.
[1856] Step 2:
[1857] Device:
[1858] Format the conversation target information entered by the user into JSON format.
[1859] Information is sent securely to the server using the HTTPS protocol.
[1860] Step 3:
[1861] server:
[1862] The received information is stored in a database.
[1863] Check the integrity of the input data and check for any abnormalities (e.g., whether name, age, gender, personality, and memories are entered correctly).
[1864] Step 4:
[1865] server:
[1866] Read the subject's information from the database.
[1867] The read information is provided as input to the AI model, and the AI model is started.
[1868] Step 5:
[1869] server:
[1870] The AI model learns the subject's personality and past memories.
[1871] Generate patterns for interaction.
[1872] Step 6:
[1873] User:
[1874] Open the chat app and tap the "Start Conversation" button.
[1875] Enter a message (e.g., "How are you doing lately?"). At the same time, the emotion engine will analyze the user's facial expressions and voice.
[1876] Step 7:
[1877] Device:
[1878] The user's message and the emotion data obtained from the emotion engine are sent to the server in JSON format.
[1879] Step 8:
[1880] server:
[1881] Receive messages and sentiment data from users.
[1882] Messages are analyzed using natural language and the emotion engine recognizes the user's emotions.
[1883] Step 9:
[1884] server:
[1885] Based on the recognized emotions and the results of natural language analysis, the AI model generates an appropriate response.
[1886] For example, if the user's emotion is recognized as "sad," the tone of the response can be adjusted to include a kind comment such as, "That must be tough. I'm rooting for you."
[1887] Step 10:
[1888] server:
[1889] The generated response is sent to the user's terminal.
[1890] Step 11:
[1891] Device:
[1892] Receives a response from the server and displays it on the chat screen.
[1893] Step 12:
[1894] User:
[1895] Continuously typing messages (e.g., "Have you traveled anywhere recently?").
[1896] Step 13:
[1897] Device:
[1898] Continue sending each user message to the server.
[1899] Step 14:
[1900] server:
[1901] Each message is received, natural language analysis is performed, and the emotion engine analyzes the user's emotions.
[1902] Based on the user's sentiment and the content of the message, the AI model generates an appropriate response (e.g., "I went to Hokkaido last week. It was so much fun!").
[1903] Step 15:
[1904] Device:
[1905] Receives each response from the server and displays it to the user.
[1906] Step 16:
[1907] User:
[1908] Tap the "End Session" button to end the chat session.
[1909] Step 17:
[1910] Device:
[1911] Sends a request to the server to end the chat session.
[1912] Step 18:
[1913] server:
[1914] Ends the current chat session based on the session ID.
[1915] Archive session data to a database or delete it as needed.
[1916] This allows users to look back on past relationships and find emotional healing by utilizing an emotion engine and AI model.
[1917] Example 2
[1918] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1919] Conventional dialogue systems have the problem that when users virtually converse with estranged friends or deceased family members, the conversations are not emotionally rich and natural. Furthermore, they lack the ability to properly recognize the user's emotions and adjust responses accordingly. This makes it difficult for users to experience emotional healing and comfort.
[1920] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input information about a conversation target (such as name, gender, age, personality, and past memories); a means for formatting the information into JSON format and transmitting it to the server via HTTPS; a means for the server to store the information in a database; a means for reading the target information from the database and training a generative AI model and an emotion engine; a means for initiating a chat session and having a dialogue between the user and the AI model; an emotion recognition means for analyzing the user's facial expressions and voice and acquiring emotion data; a means for the server to generate an appropriate response using the AI model based on the content of the dialogue and the emotion recognition results; and a means for transmitting the response to the user's terminal and displaying it on the user's chat screen. This provides a virtual dialogue experience that allows the user to reflect on past relationships and find emotional comfort and security. Recognizing emotions and adjusting responses enables a richer dialogue experience.
[1921] "User" refers to an end user who uses the dialogue system.
[1922] A "conversational target" refers to a person (friend, family, etc.) with whom the user wishes to converse.
[1923] "Information" refers to data about the person you are speaking with, such as their name, gender, age, personality, and past memories.
[1924] "JSON format" refers to a data structure in JavaScript Object Notation, a widely used format for exchanging and storing data.
[1925] The "HTTPS protocol" stands for HyperText Transfer Protocol Secure, a protocol for ensuring secure communication, and encrypts data before sending and receiving it.
[1926] "Server" refers to a computer system that processes, stores, and analyzes information submitted by users.
[1927] A "database" refers to a system that systematically organizes and stores information.
[1928] A "generative AI model" refers to a dialogue response model generated based on a machine learning algorithm, designed to enable natural dialogue with users.
[1929] An "emotion engine" refers to a software system that analyzes a user's facial expressions and voice to recognize their emotions.
[1930] A "chat session" refers to a series of conversational interactions between a user and an AI model.
[1931] "Emotion recognition means" refers to a system that provides the functionality to analyze a user's facial expressions and voice and identify their emotions.
[1932] "Natural language analysis" refers to the processes and techniques aimed at interpreting and understanding the meaning of messages from users.
[1933] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion engine, it realizes more emotionally rich dialogue. This system utilizes a generative AI model and emotion recognition function, and is implemented according to the following steps.
[1934] Hardware and software used
[1935] Hardware: The mobile devices and computers used by users
[1936] Software: Chat application, generative AI model, emotion engine, database management system (DBMS)
[1937] Overall flow
[1938] The user inputs information about the person they are talking to through a chat application. This information includes the person's name, gender, age, personality, and past memories. Once input is complete, the device formats the information into JSON format and securely transmits it to the server using the HTTPS protocol. The server stores the received information in a database and verifies the integrity of the input data. The server then reads the person's information from the database and trains the generative AI model and emotion engine. This prepares the device for a virtual conversation.
[1939] When a user starts a chat session, a message is sent from the device to the server. At the same time, the emotion engine analyzes the user's facial expressions and voice to obtain emotional data. The server performs natural language analysis and emotion recognition based on the received message and emotional data. The generative AI model then generates an appropriate response, which is sent to the user's device. This process occurs in real time, allowing the user to experience a natural, emotionally rich virtual conversation.
[1940] Specific examples
[1941] Example 1: Conversation with a friend
[1942] When a user wants to enjoy a virtual conversation with a college friend they've lost touch with, they input detailed information about the friend (such as name, gender, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message asking "How are you doing lately?" and the emotion engine determines that the response is "sad," the server will generate a response such as "I've been busy lately, but I'm glad to be able to talk to you. Is something going on?" and send it to the user. This allows for realistic, emotionally rich dialogue.
[1943] Example 2: Conversations with a deceased parent
[1944] If a user wishes to have a virtual conversation with their deceased mother, they input detailed information about her (such as her name, age, personality, and past memories). The server uses this information to train an AI model. If the user sends a message such as "Mom, do you remember me?" and the emotion engine recognizes this as "nostalgic," the server will generate a response such as "Of course! I remember you. We used to cook together a lot when you were little," and send it to the user. This allows the user to relive past memories and find emotional healing.
[1945] Prompt Sentence Examples
[1946] "I'd like to talk to my friend {Name} about {Memories}. {Name} has {Personality Traits} and we last spoke {When did we last speak}."
[1947] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1948] Step 1:
[1949] User: Enter information about the person you are talking to
[1950] A user opens a chat application and enters details about the person they are talking to (such as name, gender, age, personality, and past memories). For example, they enter the name, gender (male), age (30), personality (cheerful and sociable), and past memories (such as photos they took together on a college trip) of their friend "Friend A." They confirm the information they entered and click the "Send" button.
[1951] Input: Information about the person you are speaking to
[1952] Output: Check input information and send operation
[1953] Step 2:
[1954] Terminal: Formatting and transmitting information
[1955] The device formats the conversation target information entered by the user into JSON format, specifically creating the following data structure:
[1956] json
[1957] {
[1958] "name": "Friend A",
[1959] "gender": "male",
[1960] "age": 30,
[1961] "personality": "cheerful and sociable",
[1962] "memories": "Photos we took together on a college trip"
[1963] }
[1964] It then sends the information securely to the server using the HTTPS protocol.
[1965] Input: Information about the person you are talking to entered by the user
[1966] Output: Sends data formatted in JSON format.
[1967] Step 3:
[1968] Server: storing and verifying information
[1969] The server stores the received information in a database. Before storing it, it checks the integrity of the information, for example, whether any required fields (name, gender, age, personality, memories) are missing. If the information is determined to be correct, it records it in the database as follows:
[1970] sql
[1971] INSERT INTO conversation_data (name, gender, age, personality, memories) VALUES ('Friend A', 'Male', 30, 'Cheerful and sociable', 'Photo taken together on a college trip');
[1972] Input: Received conversation target information
[1973] Output: Records in the database and confirmation results
[1974] Step 4:
[1975] Server: Preparing the AI model
[1976] The server reads the subject's information from the database and provides it as input to the generative AI model and emotion engine. Specifically, it lists information that has not yet been learned and feeds it to the model. The AI model learns the subject's personality and past memories and generates patterns for dialogue.
[1977] Input: Subject information stored in the database
[1978] Output: Trained AI model and dialogue patterns
[1979] Step 5:
[1980] User:Start a chat session
[1981] The user opens the chat app and taps the "Start conversation" button. This is a button operation on the UI. They then type a message to their friend "Friend A" (e.g., "How are you doing lately?"). At the same time as they type the message, the emotion engine analyzes the user's facial expressions and voice in real time.
[1982] Input: Start chat session
[1983] Output: Messages and emotion data from users
[1984] Step 6:
[1985] Terminal: Sending messages and emotional data
[1986] The device sends the user's text message and the emotion data obtained from the emotion engine to the server in JSON format, generating the following data structure:
[1987] json
[1988] {
[1989] "message": "How are you doing lately?",
[1990] "emotion": "sad"
[1991] }
[1992] This data is transmitted securely using the HTTPS protocol.
[1993] Input: Message and emotion data from the user
[1994] Output: JSON data sent to the server
[1995] Step 7:
[1996] Server: Parses messages and generates responses
[1997] The server receives the user's message and emotional data and performs natural language analysis and emotion recognition. It uses natural language processing technology (e.g., NLTK or SpaCy) to analyze the message and recognize the emotional data. If the user's emotion is recognized as "sad," it adjusts the tone of the response. It uses a generative AI model to generate a response like this: "I've been busy lately, but I'm glad to talk to you. Is something going on?"
[1998] Input: Message and emotion data from the user
[1999] Output: The generated response message
[2000] Step 8:
[2001] Terminal: Display response
[2002] The device displays the response received from the server on the chat screen, and notifies the user in real time that a response has arrived.
[2003] Input: Response message from the server
[2004] Output: Display on the chat screen and notifications
[2005] (Application example 2)
[2006] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2007] In modern society, physical and emotional distance makes it difficult to reconnect with estranged friends or deceased family members. In particular, when users are experiencing emotional distress, there is a need for a way to find emotional healing through virtual conversations with these loved ones. However, conventional systems lack the means to recognize users' emotions in real time and generate appropriate responses, making it difficult to provide more natural and emotionally rich interactions.
[2008] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for saving information about the conversation target in a database, means for reading the target information from the database and training a generative AI model and an emotion analysis engine, and means for analyzing the user's facial expressions and voice and generating a response according to the emotion. This allows the user to have an emotionally rich and natural conversation through a virtual conversation with an estranged friend or a deceased family member, thereby providing emotional healing and a sense of security.
[2009] "Interviewee" refers to a person with whom you have a virtual conversation, such as a friend or family member.
[2010] A "generative AI model" is an artificial intelligence model that generates natural language based on input data.
[2011] An "emotion analysis engine" refers to software or algorithms that analyze and identify a user's emotions from their facial expressions and voice.
[2012] A "database" is a digital data storage system for storing and managing information about conversation partners and users.
[2013] A "server" is a computer system that processes information received from users and stores and manages it in a database.
[2014] "Chat session" refers to the interface through which a user and a generative AI model can converse, and the flow of that conversation.
[2015] "Natural language analysis" is the process of analyzing text entered by a user and understanding its meaning.
[2016] "Response" refers to the reply message that a generative AI model generates in response to user input.
[2017] "User" refers to the end user who uses this system to engage in conversations.
[2018] "Facial expression" refers to the element that allows a specific emotion to be read from the movement of the user's facial muscles.
[2019] "Audio" refers to sound data used to analyze emotions and content from a user's speech.
[2020] This invention is a system that provides emotional healing to users through virtual conversations with estranged friends or deceased family members, and by combining it with an emotion analysis engine, it achieves more emotionally rich dialogue. This system is implemented in the following steps.
[2021] 1. User input phase:
[2022] Users input information about the person they wish to have a virtual conversation with, including details such as name, gender, age, personality, past memories, etc. The interface for inputting this information is provided as an application on a smartphone or tablet.
[2023] 2. Data transmission phase:
[2024] The device converts the information entered by the user into JSON format and sends it to the server using the secure HTTPS protocol, which keeps the data secure.
[2025] 3. Server-side processing phase:
[2026] The server stores the received information in a database, reads the target person information from the database, and trains a generative AI model (e.g., OpenAI GPT-3) and an emotion analysis engine (e.g., Microsoft Azure Face API).
[2027] 4. Dialogue initiation phase:
[2028] When a user starts a chat session, the user's message and emotional data are sent to the server, which then performs natural language analysis and sentiment analysis. For example, a generative AI model analyzes the user's text data, and a sentiment analysis engine reads emotions from the user's facial expressions.
[2029] 5. Response generation phase:
[2030] The server uses a generative AI model based on the analysis results to generate an appropriate response, which takes into account the user's emotions. The generated response is sent to the device and displayed on the user's chat screen.
[2031] Hardware and software used
[2032] The main hardware used is smartphones and tablets. The software uses React Native as the front end, Node.js and Express as the back end, and MongoDB as the database. The sentiment analysis engine uses Microsoft Azure Face API, and the generative AI model uses OpenAI GPT-3.
[2033] Specific use cases
[2034] For example, suppose a user wants to enjoy a virtual conversation with a college friend named Ryo. The user enters Ryo's information (such as name, gender, age, personality, and memories) into the application. The server uses this information to train a generative AI model and an emotion analysis engine. If the user enters "How are you doing lately?" and the emotion analysis engine determines from the user's facial expression that they are "sad," the server will generate a response such as "I've been busy at work lately, but I'm glad I got to talk to you. Is something going on?"
[2035] Prompt Sentence Examples
[2036] I have a friend named Ryo. He's 25 years old and kind-hearted, but a bit mischievous. He remembers a trip he went on with his college club. Please respond to the following message as Ryo, expressing your feelings.
[2037] User: How are you doing lately?
[2038] Liao:
[2039] In this way, this invention provides a virtual conversational experience that allows users to reflect on past relationships and find healing and comfort. In addition, it can realize a richer experience by recognizing the user's emotions and generating more thoughtful responses.
[2040] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2041] Step 1:
[2042] The user uses a smartphone or tablet to enter information about the person they are talking to (such as name, gender, age, personality, past memories, etc.) on the application screen. The information entered by the user is converted into JSON format.
[2043] Input: User-entered information about the person you are speaking to
[2044] Output: Data converted to JSON format
[2045] Step 2:
[2046] The device securely transmits the JSON-formatted data to the server using the HTTPS protocol. Communication encryption prevents data from being intercepted by third parties.
[2047] Input: JSON format data
[2048] Output: Sending to server completed
[2049] Step 3:
[2050] The server analyzes the received JSON format data and saves it in the database. When saving, it checks the integrity of the data and eliminates invalid data.
[2051] Input: JSON format data sent to the server
[2052] Output: Data saved to database
[2053] Step 4:
[2054] The server reads the subject's information from the database, initializes the generative AI model (OpenAI GPT-3) and emotion analysis engine (Microsoft Azure Face API), and trains them.
[2055] Input: Subject information from the database
[2056] Output: Generative AI model and sentiment analysis engine initialized and trained
[2057] Step 5:
[2058] A user initiates a chat session in the application and types a message, while the device simultaneously captures the user's facial expressions with a camera and collects audio data.
[2059] Input: User-entered messages, facial expression data, and voice data
[2060] Output: Message and emotion data converted to JSON format
[2061] Step 6:
[2062] The collected messages and emotion data are then sent to the server using the HTTPS protocol, and the data is encrypted to protect the user's privacy.
[2063] Input: Message and emotion data in JSON format
[2064] Output: Sending to server completed
[2065] Step 7:
[2066] The server analyzes the received message and emotional data using natural language analysis and an emotion analysis engine. Specifically, the generative AI model analyzes the meaning of the message content, and the emotion analysis engine identifies the user's emotions from facial expressions and voice.
[2067] Input: Message and emotion data in JSON format sent to the server
[2068] Output: Analysis results
[2069] Step 8:
[2070] The server uses a generative AI model to generate an appropriate response based on the analysis results. For example, if the user's emotion is recognized as "sad," the server adjusts the tone of the response and generates words that will make the user feel lighter.
[2071] Input: Analysis results
[2072] Output: The generated response
[2073] Step 9:
[2074] The server sends the generated response to the terminal, which displays it on the chat screen. The user can then view the displayed response and enter their next message.
[2075] Input: Generated response
[2076] Output: Display on chat screen completed
[2077] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2078] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2079] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2080] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2081] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2082] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2083] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2084] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2085] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2086] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2087] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2088] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2089] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2090] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2091] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2092] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2093] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2094] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2095] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2096] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2097] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2098] The following is further disclosed regarding the above embodiment.
[2099] (Claim 1)
[2100] A means for inputting information about a conversation target;
[2101] means for transmitting the information to a server;
[2102] means for the server to store the information in a database;
[2103] A means for reading subject information from the database and training an AI model;
[2104] means for initiating a chat session and engaging in a dialogue between the user and said AI model;
[2105] means for displaying the content of the dialogue to a user;
[2106] A system including:
[2107] (Claim 2)
[2108] The system of claim 1, wherein the AI model generates dialogue patterns that reflect the personality and past memories of the person being spoken to.
[2109] (Claim 3)
[2110] The system of claim 1, wherein the server performs natural language analysis on a message from a user and generates an appropriate response.
[2111] "Example 1"
[2112] (Claim 1)
[2113] A means for inputting information about a conversation target;
[2114] means for transmitting the information to a server;
[2115] means for the server to store the information in a database;
[2116] A means for reading subject information from the database and training a generative AI model;
[2117] means for initiating a new chat session in which a user and the generative AI model interact;
[2118] A means for generating the session ID, activating the trained AI model, and associating it with the session ID;
[2119] means for displaying the content of the dialogue to a user;
[2120] A means for the server to receive a message from the user, generate an appropriate response using a generative AI model, and send it to the device;
[2121] means for displaying the generated response on a chat screen;
[2122] A system including:
[2123] (Claim 2)
[2124] The system of claim 1, wherein the generative AI model generates dialogue patterns that reflect the personality and past memories of the person being spoken to.
[2125] (Claim 3)
[2126] The system of claim 1, wherein the server performs natural language analysis on a message from a user and generates an appropriate response.
[2127] "Application Example 1"
[2128] (Claim 1)
[2129] A means for inputting information about a conversation target;
[2130] means for transmitting the information to a server;
[2131] means for the server to store the information in a database;
[2132] A means for reading subject information from the database and training an AI model;
[2133] a means for interacting with the AI model to support a physical shopping experience; and
[2134] means for displaying the content of the dialogue to a user;
[2135] A system including:
[2136] (Claim 2)
[2137] The system of claim 1, wherein the AI model generates conversation patterns that reflect the personality and past memories of the person being talked to, and provides purchasing support in a physical store.
[2138] (Claim 3)
[2139] The system of claim 1, wherein the server performs natural language analysis on a message from a user and generates an appropriate response.
[2140] "Example 2: Combining Emotion Engines"
[2141] (Claim 1)
[2142] A means for the user to input information about the person they are talking to (such as name, gender, age, personality, past memories, etc.);
[2143] means for formatting the information into JSON format and transmitting the information to a server via HTTPS protocol;
[2144] means for the server to store the information in a database;
[2145] A means for reading subject information from the database and training a generative AI model and an emotion engine;
[2146] means for initiating a chat session and engaging in a dialogue between a user and said AI model;
[2147] emotion recognition means for analyzing a user's facial expression and voice to acquire emotion data;
[2148] a means for the server to generate an appropriate response using an AI model based on the dialogue content and emotion recognition results;
[2149] means for transmitting the response to the user's terminal and displaying it on the user's chat screen;
[2150] A system including:
[2151] (Claim 2)
[2152] The system of claim 1, wherein the AI model generates dialogue patterns that reflect the personality and past memories of the person being spoken to.
[2153] (Claim 3)
[2154] 2. The system according to claim 1, wherein the server performs natural language analysis on messages and emotional data from users and generates appropriate responses using an emotional engine.
[2155] "Application example 2 when combining emotion engines"
[2156] (Claim 1)
[2157] A means for inputting information about a conversation target;
[2158] means for transmitting the information to a server;
[2159] means for the server to store the information in a database;
[2160] A means for reading subject information from the database and training a generative AI model and an emotion analysis engine;
[2161] means for initiating a chat session and engaging in a dialogue between a user and said generative AI model and sentiment analysis engine;
[2162] means for displaying the dialogue content and emotion analysis results to a user;
[2163] A system including:
[2164] (Claim 2)
[2165] The system of claim 1, wherein the generative AI model generates dialogue patterns that reflect the personality and past memories of the person being spoken to, and the emotion analysis engine analyzes the user's facial expressions and voice.
[2166] (Claim 3)
[2167] The system according to claim 1, wherein the server performs natural language analysis and sentiment analysis on messages and emotional data from users to generate appropriate responses. [Explanation of symbols]
[2168] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for inputting information about a conversation target; means for transmitting the information to a server; means for the server to store the information in a database; A means for reading subject information from the database and training an AI model; means for initiating a chat session and engaging in a dialogue between the user and said AI model; means for displaying the content of the dialogue to a user; A system including:
2. The system according to claim 1 , wherein the AI model generates dialogue patterns that reflect the personality and past memories of the person being spoken to.
3. The system according to claim 1, wherein the server performs natural language analysis on a message from a user and generates an appropriate response.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A