System
The system addresses the limitations of existing dialogue systems by incorporating a chat AI model and voice generation model to facilitate natural and personalized interactions with fictional characters, enhancing user satisfaction through individualized settings and voice recognition.
Patent Information
- Application Number
- JP2024115195
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing dialogue systems fail to provide natural and intimate interactions with fictional characters, lack the ability to save and reuse user settings, and struggle with voice message recognition and response generation, leading to unsatisfying experiences for users.
A system that includes a chat AI model for natural language processing, a voice generation model, and a database for saving user settings, enabling natural and personalized interactions by converting text to voice and managing user preferences.
Enables users to enjoy natural and intimate conversations with fictional characters, reflecting individual settings and preferences, providing a satisfying experience.
Smart Images

Figure 2026014198000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, there are many fictiosexuals, but there are limited ways to realize the interactions they desire with fictional characters. Furthermore, existing dialogue systems struggle to generate natural conversations and voice responses, preventing users from providing a satisfying experience. Furthermore, they lack the ability to save and reuse individual user settings. Given these circumstances, there is a need for a system that can provide more natural and intimate interactions with characters and reflect user settings. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means. The system includes a means for receiving an input message from a user, and a means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message. The system also includes a means for initializing a voice generation model to convert the response message into voice data, generating the voice data, and a means for sending the generated voice data to the user. The system further includes a means for verifying the user's authentication information and, if authentication is successful, displaying a dashboard, and a means for loading data of a fictional character selected by the user and performing initial setup. The system also includes a voice recognition means for converting the user's voice message into a text message, and a means for saving the conversation history and user settings in a database and loading them as needed. This system allows users to enjoy natural and intimate conversations and continue new interactions while taking individual settings into account.
[0006] "User" refers to a person who uses the system, and is an entity who accesses and operates services by being authenticated within the system.
[0007] An "input message" is text or voice data that a user sends to a system, and is information used to initiate or continue a dialogue.
[0008] A "chat AI model" is an artificial intelligence model that uses a natural language processing algorithm to generate a response message in response to a user's input message.
[0009] A "response message" is text data generated by the chat AI model, which serves as a reaction or answer to the user's input message.
[0010] A "voice generation model" is an artificial intelligence model for converting text data into voice data, and is used to provide a response message as voice.
[0011] "Audio data" refers to audio files or streams generated by a speech generation model, and is data that is played back to the user.
[0012] A "dashboard" is an interface that appears when a user logs into the system, and is a screen that provides access to various functions and settings.
[0013] "Character data" refers to information about a fictional character selected by the user, and is data necessary for initializing the chat AI model and voice generation model.
[0014] "Speech recognition means" refers to a technology for converting a user's voice message into a text message, and uses a speech recognition engine.
[0015] "Conversation history" is a record of past interactions between a user and a system, and is information stored in the form of text messages or voice data.
[0016] "User settings" refer to various parameters and preferences that the user specifies for the system, which are reflected in the tone of conversations, character reactions, etc.
[0017] "Database" means a structured data storage for storing various information managed by the system, and is used to save conversation history and user settings. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system for recreating fictional characters, and is designed to allow users to have natural conversations with the characters. Specific embodiments for implementing this system are as follows.
[0040] System Overview
[0041] This system consists of a server and a user terminal. Users can access the system through their terminal and enjoy interacting with characters. The server is responsible for the main processing and responds to user requests using a chat AI model and a voice generation model.
[0042] User authentication and character selection
[0043] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[0044] 2. Terminal: Sends the entered authentication information to the server.
[0045] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[0046] 4. User: Select the fictional character of your choice from the dashboard.
[0047] 5. Terminal: Sends the selected character information to the server.
[0048] 6. Server: Loads the chat AI model and voice generation model corresponding to the character and performs initial settings.
[0049] Conversational Interactions
[0050] 1. User: Send a text or voice message to the character from your device.
[0051] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[0052] 3. Device: Sends a text message to the server.
[0053] 4. Server: Inputs the received text message into the chat AI model and generates a response message.
[0054] 5. Chat AI model: Generates a response message and returns it to the server.
[0055] 6. Server: Input the response message into the speech generation model and generate speech data.
[0056] 7. Speech generation model: Generates speech data and returns it to the server.
[0057] 8. Server: Sends the generated voice data to the user's device.
[0058] 9. Terminal: Plays the audio data and the user hears the character's response.
[0059] Specific examples
[0060] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[0061] 1. User: Sends a text message saying, "How was your day?"
[0062] 2. Server: Receives messages and passes them to the chat AI model.
[0063] 3. Chat AI model: Analyzes the message and generates a response message such as, "Today was a really fun day."
[0064] 4. Speech generation model: Receives the response message and generates speech data.
[0065] 5. Server: Sends the audio data to the user's device.
[0066] 6. Terminal: Plays the audio data and the user hears the character's response.
[0067] Customization features
[0068] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[0069] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0070] The system allows users to enjoy natural dialogue with the characters they select, and can provide a more personalized experience through individual settings, which can satisfy the feelings of people known as fictiosexuals.
[0071] The processing flow will be explained below.
[0072] Step 1:
[0073] User: Accesses the system login screen and enters a username and password.
[0074] Step 2:
[0075] Terminal: Sends the entered authentication information to the server.
[0076] Step 3:
[0077] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[0078] Step 4:
[0079] Users: Select the fictional character of their choice from the dashboard.
[0080] Step 5:
[0081] Device: Sends the selected character information to the server.
[0082] Step 6:
[0083] Server: Loads the chat AI model and voice generation model corresponding to the characters and performs initial settings.
[0084] Step 7:
[0085] User: Send a text or voice message from your device to the character.
[0086] Step 8:
[0087] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0088] Step 9:
[0089] Terminal: Sends a text message to the server.
[0090] Step 10:
[0091] Server: Inputs the received text message into the chat AI model and generates a response message.
[0092] Step 11:
[0093] Chat AI model: Analyzes messages and generates appropriate response messages.
[0094] Step 12:
[0095] Server: Input the response message into the speech generation model and generate speech data.
[0096] Step 13:
[0097] Speech generation model: Analyzes text data and generates speech data.
[0098] Step 14:
[0099] Server: Sends the generated audio data to the user's device.
[0100] Step 15:
[0101] Terminal: Plays the audio data and the user hears the character's response.
[0102] Step 16:
[0103] User: Set preferences to customize the tone of the dialogue and character reactions.
[0104] Step 17:
[0105] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0106] This series of steps allows users to enjoy natural interactions with their chosen fictional characters.
[0107] Example 1
[0108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0109] With conventional dialogue systems, it has been difficult to achieve natural and personalized dialogue with fictional characters. It has also been difficult to properly reflect the user's ratings and preferences and respond individually. This has resulted in inconsistent dialogue experiences with the selected characters, leading to low satisfaction. Furthermore, the process of voice message recognition and response generation is complex, making system setup and use time-consuming.
[0110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0111] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for sending the generated voice data to the user, means for comparing authentication information entered by the user with a database and displaying an access portal if authentication is successful, means for loading data of a virtual character selected by the user and performing initial settings, means for speech recognition to convert voice messages into text, and means for saving conversation history and user settings in a database and loading them as needed, thereby enabling natural and personalized interactions with fictional characters.
[0112] The "means for receiving an input message from a user" refers to a device or program that has the function of transmitting a text message or voice message input by a user to a server.
[0113] "Means for initializing a chat AI model using a natural language processing algorithm and generating a response message" refers to a device or program that has the function of initiating the operation of a chat AI model using natural language processing technology and generating an appropriate response to a message from a user.
[0114] The "means for initializing a speech generation model and generating speech data" refers to a device or program having the function of operating a speech generation model for converting a text message into speech data and converting the generated text response into speech data.
[0115] The "means for transmitting generated voice data to a user" refers to a device or program having the function of transmitting voice data generated by the server to a user's terminal.
[0116] "Means for checking the authentication information entered by the user against a database and displaying an access portal if authentication is successful" refers to a device or program that has the function of checking the authentication information entered by the user (such as a user name and password) against a database to perform authentication, and displaying the main screen if authentication is successful.
[0117] "Means for loading the data of a virtual character selected by the user and performing initial settings" refers to a device or program that has the function of loading the data of a character selected by the user and setting the corresponding chat AI model and voice generation model.
[0118] The "voice recognition means for converting a voice message into text" is a device or program that has the function of converting a voice message uttered by a user into text format.
[0119] "Means for saving conversation history and user settings in a database and loading them as needed" refers to a device or program that has the function of recording the conversation history and user settings in a database and referencing this information during subsequent conversations.
[0120] This invention relates to a system that allows users to interact with fictional characters. The system consists of a server and a user terminal, and utilizes a chat AI model and a voice generation model to realize natural and personalized interactions.
[0121] System configuration
[0122] The system mainly consists of the following components:
[0123] 1. Server: A high-performance computer is used to process the chat AI model and voice generation model.
[0124] 2. Terminal: The device through which the user accesses the service, including smartphones and PCs.
[0125] 3. Database: Stores authentication information, character settings, conversation history, and user preferences.
[0126] Hardware and software used
[0127] Hardware: High-performance servers (e.g., cloud service provider servers), user devices (smartphones, PCs)
[0128] Software: Database (e.g., MySQL), chat AI model (e.g., GPT-3), speech generation model (e.g., Google Cloud Text-to-Speech), speech recognition method (e.g., Google Cloud Speech-to-Text)
[0129] Explanation of program processing
[0130] User authentication
[0131] The user accesses the system's login screen from a terminal and enters their username and password. The terminal encrypts this authentication information and sends it to the server, which checks it against a database and, if authentication is successful, displays the access portal. This process uses a MySQL database.
[0132] Character Selection
[0133] The user selects the character they want to interact with from the access portal. The device sends the character information to the server, which then loads the chat AI model (e.g., GPT-3) and voice generation model (e.g., Google Cloud Text-to-Speech) corresponding to the selected character and performs initial setup.
[0134] Conversational Interactions
[0135] Users send text or voice messages from their devices. When a voice message is sent, the device converts the speech to text using Google Cloud Speech-to-Text. The device then sends the text message to the server. The server inputs the message into a chat AI model and generates a reply message. The generated reply message is then input into a voice generation model, which generates voice data. The server then sends this voice data to the user's device, which then plays it back.
[0136] Customization features
[0137] Users can customize the tone of the conversation and the character's reactions in the settings screen, and this information is stored in a database on the server and applied to future interactions.
[0138] Specific examples
[0139] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[0140] User: Texts "How was your day?"
[0141] Terminal: Sends a text message to the server.
[0142] Server: Passes the message to the chat AI model (GPT-3) and generates a response message.
[0143] Chat AI model: Generates the response "Today was a really fun day."
[0144] Speech generation model: Converts the response message into speech data.
[0145] Server: Sends the audio data to the user's device.
[0146] Terminal: Plays the audio data and the user hears the character's response.
[0147] Prompt Sentence Examples
[0148] "Ask Character A, 'How was your day?'"
[0149] The invention allows users to enjoy natural and personalized interactions with fictional characters.
[0150] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0151] Step 1:
[0152] User: Access the device's login screen and enter your username and password.
[0153] Input: Username and Password.
[0154] Output: The authentication information is sent to the device.
[0155] Step 2:
[0156] On the device: The authentication information entered by the user is encrypted and sent to the server.
[0157] Input: Encrypted credentials.
[0158] Output: Authentication information sent to the server.
[0159] Step 3:
[0160] Server: Compares the user information stored in the database, and if authentication is successful, displays the access portal; if authentication fails, generates and returns an error message.
[0161] Input: Encrypted credentials.
[0162] Output: The authentication result (success or failure).
[0163] Step 4:
[0164] User: Select the desired character from the character list displayed on the dashboard.
[0165] Input: The ID of the character selected by the user.
[0166] Output: Character selection information is sent to the device.
[0167] Step 5:
[0168] Device: Sends the selected character information to the server.
[0169] Input: Character ID.
[0170] Output: Character information is sent to the server.
[0171] Step 6:
[0172] Server: Loads the chat AI model and voice generation model corresponding to the selected character and performs initial settings.
[0173] Input: Character ID.
[0174] Output: The chat AI model and speech generation model are initialized.
[0175] Step 7:
[0176] User: Send a text or voice message to the character saying, "How was your day?"
[0177] Input: The user's text or voice message.
[0178] Output: The message is sent to the terminal.
[0179] Step 8:
[0180] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0181] Input: The user's voice message.
[0182] Output: A text message.
[0183] Step 9:
[0184] Terminal: Sends a text message to the server.
[0185] Input: A text message.
[0186] Output: A message is sent to the server.
[0187] Step 10:
[0188] Server: Inputs the received text message into the chat AI model and generates a response message.
[0189] Input: A text message.
[0190] Output: The response message.
[0191] Step 11:
[0192] Chat AI model: Analyzes messages and generates a response message such as "I had a great day today."
[0193] Input: The user's text message.
[0194] Output: The response message.
[0195] Step 12:
[0196] Server: Input the response message into the speech generation model and generate speech data.
[0197] Input: The response message.
[0198] Output: Audio data.
[0199] Step 13:
[0200] Speech generation model: Receives the response message and generates speech data.
[0201] Input: The response message.
[0202] Output: Audio data.
[0203] Step 14:
[0204] Server: Sends the generated audio data to the user's device.
[0205] Input: Audio data.
[0206] Output: The audio data is sent to the device.
[0207] Step 15:
[0208] Terminal: Plays the audio data and the user hears the character's response.
[0209] Input: Audio data.
[0210] Output: Playback of audio data.
[0211] (Application example 1)
[0212] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0213] In virtual environments, users are expected to have a natural experience through interactions with fictional characters. However, conventional systems have made it difficult for users to interact with characters in real time within a virtual environment and receive product explanations and guidance, making it difficult to provide a truly immersive experience. The present invention aims to solve these problems by providing a system that allows users to interact with fictional characters in a virtual environment in a natural and interactive manner, and to receive product explanations and guidance.
[0214] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0215] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for transmitting the generated voice data to the user, and means for reproducing a fictional character in the virtual environment based on a selection by the user and providing guidance or product explanations, thereby enabling the user to have natural conversations with the fictional character in the virtual environment in real time and receive product explanations or guidance.
[0216] "Means for receiving input messages from a user" means a mechanism by which the system receives messages input by a user when the user interacts with a fictional character within a virtual environment.
[0217] A "chat AI model" is an artificial intelligence model built on a natural language processing algorithm used to generate natural response messages in response to user input messages.
[0218] A "voice generation model" is a model used to convert a text response message generated by a chat AI model into voice data.
[0219] A "virtual environment" is a computer-generated virtual space experienced by a user, for example, using a VR headset or smart device.
[0220] A "fictional character" is a fictional character that appears in fictional works such as movies, anime, and games, and whose role is to provide an experience through interactions with users.
[0221] "User authentication" is the process of identifying and verifying a user's identity when accessing a system.
[0222] The "dashboard" is the interface that users can access after authentication, and is the main screen for selecting fictional characters.
[0223] "Real-time response" is a function that generates a response immediately to user input, allowing the character to respond immediately.
[0224] A "database" is a system that stores information such as user settings and conversation history and can read that information as needed.
[0225] "Guide and product information" refers to an explanation or introduction given to a user by a fictional character, providing product details and purchasing advice.
[0226] A specific embodiment of the present invention is to build a system that allows users to naturally interact with fictional characters in a virtual environment. This system is composed of a user terminal and a cloud server.
[0227] System Configuration
[0228] Hardware and Software
[0229] User devices: smartphones (iOS or Android), VR headsets (e.g., Oculus Quest 2), smart glasses.
[0230] Server: Cloud service (e.g. AWS EC2).
[0231] Database: A cloud database service (e.g. AWS RDS).
[0232] AI models: Chat AI models (e.g., OpenAI GPT-4), speech generation models (e.g., Google Wavenet).
[0233] Program processing
[0234] 1. User authentication:
[0235] When a user logs in to the app, the device sends authentication information to the server.
[0236] The server checks the authentication information against the database, and if authentication is successful, displays the dashboard on the terminal.
[0237] 2. Character Selection:
[0238] Users select a fictional character from the dashboard.
[0239] The device sends the selected character information to the server, and the server loads the chat AI model and voice generation model for the character and performs initial settings.
[0240] 3. Conversational Interactions:
[0241] When the user types "Hello, what do you recommend today?", the device sends this message to the server.
[0242] The server passes the input message to the chat AI model to generate a response message.
[0243] The chat AI model generates a response message saying, "Hello! This month's recommendation is new earphones," and returns it to the server.
[0244] The server sends a response message to the speech generation model to generate speech data.
[0245] The generated audio data is returned to the terminal, which then plays the audio.
[0246] 4. Enhanced user experience:
[0247] Users can enjoy a virtual experience in which fictional characters act as guides, showing them around the store and explaining products.
[0248] Conversation history and user preferences are stored in a database for later reuse.
[0249] Specific examples
[0250] For example, if the user selects "Fictional Character B":
[0251] 1. Put on the VR headset.
[0252] 2. A user sends a text message asking, "What are your recommendations in the store?"
[0253] 3. The chat AI model responds, "The recommended product is the newly released smartwatch."
[0254] 4. The speech generation model converts the response into speech.
[0255] 5. Playback on the user's device.
[0256] In this way, it is possible to provide users with an engaging experience by realizing natural interactions and real-time responses with fictional characters within a virtual environment.
[0257] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0258] Step 1:
[0259] User authentication
[0260] The user logs in to the app and enters their user ID and password.
[0261] The terminal transmits the input authentication information to the server.
[0262] The server checks the authentication information against the database, and if authentication is successful, generates dashboard information along with an authentication success message and sends it to the terminal. If authentication fails, it returns an error message.
[0263] The device displays the received dashboard information to the user and waits for the next action.
[0264] Step 2:
[0265] Character Selection
[0266] Users select the fictional character they want from the dashboard.
[0267] The terminal transmits the selected character information to the server and displays a confirmation message.
[0268] The server loads the chat AI model and voice generation model for the character, performs initial setup, and returns a message to the device indicating that the setup is complete.
[0269] The terminal will display an initialization complete message to the user and wait for further action.
[0270] Step 3:
[0271] Interaction Start
[0272] The user types a text message (e.g., "Hello, what do you recommend today?") or a voice message to the selected character.
[0273] If the message is a voice message, the terminal performs voice recognition processing to convert it into text, and then sends the text message to the server.
[0274] The server inputs the received text message into the chat AI model and generates a response message.
[0275] The chat AI model generates a response message such as "Hello! This month's recommendation is new earphones" and returns it to the server.
[0276] Step 4:
[0277] Generating a response message
[0278] The server sends the response message received from the chat AI model to the voice generation model, which generates voice data.
[0279] The speech generation model analyzes the text message and generates corresponding speech data, which is returned to the server.
[0280] Step 5:
[0281] Sending and playing reply messages
[0282] The server transmits the generated voice data to the terminal.
[0283] The device plays back the received audio data, allowing the user to hear the fictional character's response.
[0284] Along with the response message, it displays UI elements to determine the interaction status and next action.
[0285] Step 6:
[0286] Save conversation history and settings
[0287] The server stores conversation history and user settings (e.g. character customization) in a database.
[0288] The database stores the necessary information to allow users to access their conversation history and preferences the next time they log in.
[0289] This series of processes allows users to have natural conversations with fictional characters in a virtual environment, receive responses in real time, and enjoy product guides and explanations in a virtual store.
[0290] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0291] This system recreates fictional characters and allows natural dialogue with users. It also incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly, providing more personalized interactions. The system consists of a server and a user terminal.
[0292] System Overview
[0293] The system's server handles the main processing and responds to user requests using a chat AI model, a voice generation model, and an emotion engine. Users can access the system via their devices and enjoy interacting with the characters.
[0294] User authentication and character selection
[0295] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[0296] 2. Terminal: Sends the entered authentication information to the server.
[0297] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[0298] 4. User: Select the fictional character of your choice from the dashboard.
[0299] 5. Terminal: Sends the selected character information to the server.
[0300] 6. Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[0301] Conversational Interactions
[0302] 1. User: Send a text or voice message to the character from your device.
[0303] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[0304] 3. Device: Sends a text message to the server.
[0305] 4. Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[0306] 5. Emotion Engine: Recognizes user emotions from text messages and generates emotional status.
[0307] 6. Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[0308] 7. Chat AI model: Generates a response message and returns it to the server.
[0309] 8. Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[0310] 9. Speech generation model: Generates speech data based on text data and emotional status.
[0311] 10. Server: Sends the generated voice data to the user's device.
[0312] 11. Terminal: Plays the audio data and the user hears the character's response.
[0313] Specific examples
[0314] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[0315] 1. User: Texts "I'm so tired today."
[0316] 2. Server: Receives the message and passes it to the emotion engine.
[0317] 3. Emotion engine: Analyzes the message and generates an emotional status that identifies the user as tired.
[0318] 4. Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been tough, take a break."
[0319] 5. Speech generation model: Receives a response message and generates speech data in a gentle tone according to the emotional status.
[0320] 6. Server: Sends the audio data to the user's device.
[0321] 7. Terminal: Play the audio data and the character responds in a gentle tone, saying, "That must have been hard. Take a break."
[0322] Customization features
[0323] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[0324] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0325] This system allows users to enjoy natural and emotional interactions with their chosen fictional characters, and individual settings can provide a more personalized experience, allowing for deeper emotional fulfillment for people known as fictiosexuals.
[0326] The processing flow will be explained below.
[0327] This invention is a system that recreates fictional characters and enables natural dialogue with users, and by incorporating an emotion engine that recognizes the user's emotions and adjusts the response content, it provides more personalized interactions. The specific processing flow of the system is explained below in steps.
[0328] User authentication and character selection
[0329] Step 1:
[0330] User: Accesses the system login screen from a terminal and enters a username and password.
[0331] Step 2:
[0332] Terminal: Sends the entered authentication information to the server.
[0333] Step 3:
[0334] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[0335] Step 4:
[0336] Users: Select the fictional character of their choice from the dashboard.
[0337] Step 5:
[0338] Device: Sends the selected character information to the server.
[0339] Step 6:
[0340] Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[0341] Conversational Interactions
[0342] Step 7:
[0343] User: Send a text or voice message from your device to the character.
[0344] Step 8:
[0345] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0346] Step 9:
[0347] Terminal: Sends a text message to the server.
[0348] Step 10:
[0349] Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[0350] Step 11:
[0351] Emotion Engine: Recognizes user emotions from text messages and generates emotional statuses.
[0352] Step 12:
[0353] Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[0354] Step 13:
[0355] Chat AI model: Analyzes messages and generates appropriate response messages.
[0356] Step 14:
[0357] Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[0358] Step 15:
[0359] Speech generation model: Generates speech data based on text data and emotional status.
[0360] Step 16:
[0361] Server: Sends the generated audio data to the user's device.
[0362] Step 17:
[0363] Terminal: Plays the audio data and the user hears the character's response.
[0364] Specific examples
[0365] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[0366] Step 1:
[0367] User: Texts "I'm so tired today."
[0368] Step 2:
[0369] Server: Receives messages and passes them to the emotion engine.
[0370] Step 3:
[0371] Emotion engine: Analyzes messages and generates an emotional status that identifies the user as tired.
[0372] Step 4:
[0373] Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been hard, take a break."
[0374] Step 5:
[0375] Speech generation model: Receives a response message and generates voice data in a gentle tone according to the emotional status.
[0376] Step 6:
[0377] Server: Sends the audio data to the user's device.
[0378] Step 7:
[0379] Device: Plays the audio data and the character responds in a gentle tone, "That must have been hard, take a break."
[0380] Customization features
[0381] Step 1:
[0382] User: You can set preferences to customize the tone of the dialogue and character reactions.
[0383] Step 2:
[0384] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0385] This allows the system to realize natural and emotional interactions with fictional characters selected by the user, making it possible to more deeply satisfy the emotions of people known as fictiosexuals.
[0386] Example 2
[0387] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0388] Conventional dialogue systems have difficulty in engaging in natural conversations with fictional characters, limiting their ability to provide personalized responses based on the user's emotions. Furthermore, the accuracy of speech recognition and speech synthesis is low, making it difficult to generate speech that takes emotions into account. Furthermore, systems lack a mechanism for centrally managing user authentication information and settings and saving dialogue history to provide a better user experience. Therefore, a new dialogue system is needed to improve user satisfaction.
[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0390] In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue generation model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice synthesis model to convert the response message into voice data and generating voice data, means for utilizing an emotion analysis engine to analyze the user's emotion from the input message, means for generating a response message taking into account the emotional status obtained by the emotion analysis engine, means for utilizing the voice synthesis model to generate voice data with a tone corresponding to the emotional status, and means for transmitting the generated voice data to the user. This makes it possible to generate a response message and voice data that take into account the user's emotion and provide a natural and personalized dialogue with a fictional character.
[0391] "User" means any individual or entity that accesses the System and interacts with fictional characters.
[0392] "Input Message" refers to text or voice information sent by a user to a system.
[0393] "Natural language processing algorithms" refer to technologies for understanding a user's input message and generating an appropriate response.
[0394] "Dialogue generation model" refers to a machine learning model that uses a natural language processing algorithm to generate a response message to a user's input message.
[0395] "Response message" refers to text information that the dialogue generation model generates in response to a user's input message.
[0396] "Speech synthesis model" refers to the technology for converting text response messages into speech data.
[0397] "Audio Data" refers to an audio file or audio stream generated by a speech synthesis model.
[0398] "Sentiment analysis engine" refers to technology for analyzing emotions from a user's input message and generating an emotional status.
[0399] "Emotional status" refers to information that expresses a user's emotions numerically or categorically, as determined by the emotion analysis engine.
[0400] "Authentication Information" refers to information such as username and password that a user uses to log into a system.
[0401] "Dashboard" refers to the interface that allows users to select characters and change settings.
[0402] "Fictional Character" refers to any imaginary or fictional person or entity that is recreated within the system.
[0403] "Voice recognition means" refers to technology used to convert a voice message sent by a user into a text message.
[0404] "Conversation history" refers to a record of the dialogue between a user and a character.
[0405] "Database" refers to a management system for systematically storing information necessary for system operation.
[0406] This invention is a system that allows users to have natural and emotional conversations with fictional characters. This system analyzes the user's emotions and appropriately adjusts the content and tone of the conversation to provide a more personalized conversation. Specifically, the system uses the following hardware and software to process data and perform calculations.
[0407] Hardware and software used
[0408] server
[0409] Database Management System (DBMS): Stores user authentication information and configuration data, typically using MySQL or PostgreSQL.
[0410] Dialogue generation model: Generates dialogue. As an example, GPT-3, a machine learning model, is used.
[0411] Speech synthesis model: Converts the response message into audio data. For example, speech synthesis APIs such as Google Cloud Text-to-Speech and IBM Watson Text to Speech are used.
[0412] Sentiment analysis engine: Recognizes user emotions. For example, Microsoft Azure's Sentiment Analysis is used.
[0413] Terminal
[0414] User Interface (UI): A screen where users can interact with fictional characters. Most commonly, it is a web app that is made up of HTML / CSS and JavaScript.
[0415] Speech recognition method: Converts voice input into text. Google Cloud Speech-to-Text and IBM Watson Speech to Text are used.
[0416] User authentication and character selection
[0417] The system begins when a user logs in using their authentication information. The user accesses the system's login screen from their terminal and enters their username and password. The input information is sent from the terminal to the server, which then checks it against a database to authenticate the user.
[0418] After successful authentication, the user is redirected to a dashboard where they can select their desired fictional character. The selected character information is sent from the device to the server, which then loads and initializes the character's corresponding dialogue generation model, speech synthesis model, and sentiment analysis engine.
[0419] Conversational Interactions
[0420] When a user starts a dialogue with a character, an input message is sent from the terminal to the server. If this input message is a voice message, the terminal converts it into text using a voice recognition means.
[0421] The received text message is processed by the server and passed to the sentiment analysis engine, which analyzes the message and generates the user's emotional status. The emotional status is input into the dialogue generation model, and a response message is generated that takes the user's emotions into account.
[0422] The generated response message and emotional status are passed to a speech synthesis model to generate speech data, including tones, which is then sent from the server to the user's device and played back on the device.
[0423] Customization features
[0424] The system provides users with settings to customize the tone of dialogue and character reactions. Users can adjust various settings by manipulating sliders and checkboxes on the settings page. These settings are saved in a database by the server and applied the next time a dialogue is held.
[0425] Specific examples
[0426] For example, if a user sends a message to "Fictional Character A" saying "I'm really tired today," the system operates as follows: First, the user sends a text message saying "I'm really tired today." The server receives this message and passes it to the emotion analysis engine. The emotion analysis engine recognizes that the user is tired and generates an emotion status.
[0427] Next, the server inputs the emotional status into a dialogue generation model and generates a response message that matches the user's emotion, "That must have been hard. Take a break." The speech synthesis model receives this response message and generates voice data in a gentle tone according to the emotional status. Finally, the server sends the voice data to the user's device, which plays it back, with the character responding in a gentle tone, "That must have been hard. Take a break."
[0428] Prompt Sentence Examples
[0429] If a user sends a message to fictional character A saying, "I'm really tired today," please explain in detail how the emotion engine generates the emotion status and what response the dialogue generation model makes.
[0430] As described above, the present invention is a system that allows users to enjoy natural and emotional interactions with fictional characters of their choice, providing a more personalized experience through individual settings.
[0431] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0432] Divide the program's processing flow into processing steps
[0433] Step 1:
[0434] Step 2:
[0435] Step 3:
[0436] ...
[0437] Specific explanations for each step
[0438] Step 1: User authentication and login
[0439] Input: Username, Password
[0440] Output: Authentication result (success or failure), dashboard screen or error message
[0441] Operation:
[0442] 1. User: Enter your username and password on the device login screen.
[0443] 2. Terminal: Sends the input information to the server as a POST request.
[0444] 3. Server: Query the database to verify the authentication information. If authentication is successful, return the dashboard screen; if authentication is unsuccessful, return an error message.
[0445] Step 2: Choose your character
[0446] Input: Character ID
[0447] Output: Character data, initialization result
[0448] Operation:
[0449] 1. User: Select the fictional character of your choice from the dashboard.
[0450] 2. Device: Send the selected character ID to the server as an AJAX request.
[0451] 3. Server: Loads the dialogue generation model, speech synthesis model, and emotion analysis engine corresponding to the character and performs initial configuration.
[0452] Step 3: Send the message
[0453] Input: Text or voice message
[0454] Output: Text message (in case of voice)
[0455] Operation:
[0456] 1. User: Send a text or voice message from your device to the character.
[0457] 2. Device: When a voice message is sent, a speech recognition tool is used to convert the voice into text, the audio file is sent to a cloud API, and the text is returned.
[0458] Step 4: Sentiment Analysis
[0459] Input: Text message
[0460] Output: Emotional status
[0461] Operation:
[0462] 1. Terminal: Sends the converted text message to the server.
[0463] 2. Server: Inputs the received text message into the sentiment analysis engine to analyze the user's sentiment. The sentiment engine analyzes the message and generates a sentiment score.
[0464] Step 5: Generate a response message
[0465] Input: Text message, emotional status
[0466] Output: Response message
[0467] Operation:
[0468] 1. Server: Inputs the emotional status into the dialogue generation model and generates a response message that takes the user's emotions into account. The generative AI model analyzes the text message and emotional status and generates an appropriate response.
[0469] Step 6: Generate audio data
[0470] Input: Response message, emotional status
[0471] Output: Audio data
[0472] Operation:
[0473] 1. Server: Input the response message and emotional status into the speech synthesis model and generate speech data with a tone corresponding to the emotion. The speech synthesis API generates an audio file with a tone based on the emotion.
[0474] Step 7: Playing back audio data
[0475] Input: Audio data
[0476] Output: Play audio
[0477] Operation:
[0478] 1. Server: Sends the generated voice data to the user's device.
[0479] 2. Device: The audio data is played and the user hears the character's response. The device's audio player plays the audio data.
[0480] This is the flow of the program's processing. The specific operations, inputs, and outputs at each step are explained in detail, and it contains all the elements necessary for users to enjoy natural and emotional interactions with fictional characters.
[0481] (Application example 2)
[0482] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0483] Conventional chat systems and voice dialogue systems have difficulty responding to user emotions during interactions, resulting in a limited user experience. In particular, when dealing with customers in physical stores, flexible responses based on customer emotions are required, but the means to achieve this are not fully developed. As a result, there is a risk of customer satisfaction decreasing.
[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue AI model using a natural language processing algorithm for responding to the input message and generating a response message, means for initializing a voice generation algorithm for converting the response message into voice data and generating voice data, means for adding emotion to the response message using an emotion analysis engine that analyzes the user's emotion, and means for transmitting the generated voice data to the user. This enables natural and personalized dialogue that takes the user's emotion into consideration.
[0485] "Means for receiving input messages from a user" means a system or function for obtaining text or voice messages sent by a user from a terminal.
[0486] A "dialogue AI model using natural language processing algorithms" is an algorithm or model that analyzes a user's input message and generates a response accordingly.
[0487] "Speech generation algorithm" refers to an algorithm or technique for converting a text response message into voice data.
[0488] An "emotion analysis engine" is an analysis system that identifies emotions from the user's input message and reflects those emotions in the response message.
[0489] "Means for verifying user authentication information" refers to a system or function that checks authentication information such as a user's ID and password to determine whether the user is legitimate.
[0490] The "means for displaying a dashboard" refers to a system or function for providing an operation screen that is displayed after a user logs in to the system.
[0491] A "means for loading fictional character data" is a system or function that loads data or settings related to a specific fictional character into memory and makes them available for use.
[0492] "Means for analyzing the user's emotional status and reflecting it in the initial settings" refers to a system or function for analyzing the user's emotions and reflecting the results in the character's settings and behavior.
[0493] "Speech recognition means" refers to technology or systems that convert the user's speech into text.
[0494] "Means for storing conversation history and user settings in a database" refers to a system or function that records the content of conversations with users and user settings (such as customization information) in a database, allowing them to be searched or referenced when necessary.
[0495] A "means for customizing fictional character responses" is a system or feature that allows a fictional character's words and actions to be flexibly changed based on the user's settings and choices.
[0496] In this invention, a natural, emotion-conscious dialogue between a user and a fictional character can be realized using the following system configuration and processing procedure.
[0497] Overall system configuration
[0498] The server contains the following main modules:
[0499] 1. Input message receiving module
[0500] Receive text and voice messages from user devices.
[0501] 2. Dialogue AI Model
[0502] It uses natural language processing algorithms to respond to input messages from users.
[0503] 3. Speech Generation Algorithm
[0504] The response message is converted into voice data.
[0505] 4. Sentiment Analysis Engine
[0506] Analyze the user's emotions from the input message and reflect them in the response message.
[0507] 5. User Authentication Module
[0508] Verify the user's credentials and display the dashboard if authentication is successful.
[0509] 6. Character Data Load Module
[0510] Loads a fictional character of the user's choice and performs initial setup.
[0511] 7. Speech Recognition Module
[0512] Convert voice messages to text.
[0513] 8. Database Management Module
[0514] Conversation history and user preferences are stored in a database and loaded as needed.
[0515] 9. Character Reaction Customization Module
[0516] Customize character reactions based on user choices and settings.
[0517] Hardware and Software
[0518] The hardware and software used are as follows:
[0519] Server hardware: A server with a powerful processor and sufficient memory and storage. For example, you can use cloud services such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0520] Software: Generative AI models such as GPT-3 are used for natural language processing algorithms. WaveNet and Tacotron are used for speech generation algorithms. Sentiment analysis libraries (Python text2emotion, transformers, etc.) are used for sentiment analysis.
[0521] Specific explanation of the process
[0522] server
[0523] 1. The input message receiving module receives a message from the user, which can be text or voice.
[0524] 2. The speech recognition module converts the voice message into text.
[0525] 3. The converted text is input into a sentiment analysis engine to analyze the user's sentiment.
[0526] 4. The analyzed emotional status is input into a dialogue AI model to generate a response message that takes emotions into account.
[0527] 5. The response message is converted into voice data by a voice generation algorithm.
[0528] 6. The generated voice data is sent to the user terminal.
[0529] 7. The database management module stores conversation history and user settings for future use.
[0530] 8. Character reaction customization module modifies character reactions based on user settings.
[0531] Examples and prompts
[0532] For example, if a customer asks "What products do you recommend?" in a physical store, the specific processing is as follows:
[0533] User: Ask "What products do you recommend?"
[0534] Server: The sentiment analysis engine analyzes the sentiment of the question "What product do you recommend?", and the dialogue AI model generates a response message.
[0535] The voice generation algorithm responds, "Today's recommendation is fresh apples," and the generated voice data is sent to the user's device.
[0536] Example prompt sentence:
[0537] "Please enter a user message: What products do you recommend?"
[0538] "Please perform sentiment analysis on this message and generate an appropriate response."
[0539] In this way, a natural and emotional interaction with the user can be achieved.
[0540] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0541] Step 1:
[0542] The user inputs a voice message or text message from a device such as a smartphone or tablet, and the input message is sent from the device to the server.
[0543] Step 2:
[0544] The server invokes a speech recognition module to analyze the received message. When a voice message is sent, the speech recognition module converts the voice data into text data. The text data is used as input and the speech recognition result is obtained as output.
[0545] Step 3:
[0546] The server inputs the text message into the sentiment analysis engine. The sentiment analysis engine analyzes the emotional status from the user's input message. In this step, the sentiment analysis engine analyzes the text data and extracts the emotional status from the user's text message. The input is the text data and the output is the emotional status.
[0547] Step 4:
[0548] The server inputs the analyzed emotional status into a dialogue AI model to generate a response message that takes the user's emotions into account. The dialogue AI model is based on a generative AI model, which receives the emotional status and text message as input and generates a response message as output. This process results in an appropriate response that corresponds to the user's input and emotions.
[0549] Step 5:
[0550] The server sends the generated response message to a voice generation algorithm to generate voice data. The voice generation algorithm receives the response message and the emotional status as input and outputs the voice data in a tone that reflects the emotion.
[0551] Step 6:
[0552] The server transmits the generated voice data to the user's device, which then plays the received voice data and provides the character's response to the user, allowing the user to experience voice responses that reflect their emotions.
[0553] Step 7:
[0554] The server stores the conversation history and user settings in a database. The database management module uses the conversation data and settings information as input and obtains the saved data as output. This data can be used as a reference for future conversations.
[0555] Step 8:
[0556] The user customizes the character's reactions as needed. The server's character reaction customization module receives the user's settings and generates the character's reactions based on them. In this step, the user's input settings are processed and the customized reactions are output.
[0557] The above are the processing steps of the system that realizes this invention, which allows the user to have a natural conversation experience that takes emotions into consideration.
[0558] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0559] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0560] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0561] [Second embodiment]
[0562] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0563] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0564] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0565] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0566] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0567] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0568] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0569] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0570] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0571] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0572] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0573] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0574] The present invention is a system for recreating fictional characters, and is designed to allow users to have natural conversations with the characters. Specific embodiments for implementing this system are as follows.
[0575] System Overview
[0576] This system consists of a server and a user terminal. Users can access the system through their terminal and enjoy interacting with characters. The server is responsible for the main processing and responds to user requests using a chat AI model and a voice generation model.
[0577] User authentication and character selection
[0578] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[0579] 2. Terminal: Sends the entered authentication information to the server.
[0580] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[0581] 4. User: Select the fictional character of your choice from the dashboard.
[0582] 5. Terminal: Sends the selected character information to the server.
[0583] 6. Server: Loads the chat AI model and voice generation model corresponding to the character and performs initial settings.
[0584] Conversational Interactions
[0585] 1. User: Send a text or voice message to the character from your device.
[0586] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[0587] 3. Device: Sends a text message to the server.
[0588] 4. Server: Inputs the received text message into the chat AI model and generates a response message.
[0589] 5. Chat AI model: Generates a response message and returns it to the server.
[0590] 6. Server: Input the response message into the speech generation model and generate speech data.
[0591] 7. Speech generation model: Generates speech data and returns it to the server.
[0592] 8. Server: Sends the generated voice data to the user's device.
[0593] 9. Terminal: Plays the audio data and the user hears the character's response.
[0594] Specific examples
[0595] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[0596] 1. User: Sends a text message saying, "How was your day?"
[0597] 2. Server: Receives messages and passes them to the chat AI model.
[0598] 3. Chat AI model: Analyzes the message and generates a response message such as, "Today was a really fun day."
[0599] 4. Speech generation model: Receives the response message and generates speech data.
[0600] 5. Server: Sends the audio data to the user's device.
[0601] 6. Terminal: Plays the audio data and the user hears the character's response.
[0602] Customization features
[0603] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[0604] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0605] The system allows users to enjoy natural dialogue with the characters they select, and can provide a more personalized experience through individual settings, which can satisfy the feelings of people known as fictiosexuals.
[0606] The processing flow will be explained below.
[0607] Step 1:
[0608] User: Accesses the system login screen and enters a username and password.
[0609] Step 2:
[0610] Terminal: Sends the entered authentication information to the server.
[0611] Step 3:
[0612] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[0613] Step 4:
[0614] Users: Select the fictional character of their choice from the dashboard.
[0615] Step 5:
[0616] Device: Sends the selected character information to the server.
[0617] Step 6:
[0618] Server: Loads the chat AI model and voice generation model corresponding to the characters and performs initial settings.
[0619] Step 7:
[0620] User: Send a text or voice message from your device to the character.
[0621] Step 8:
[0622] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0623] Step 9:
[0624] Terminal: Sends a text message to the server.
[0625] Step 10:
[0626] Server: Inputs the received text message into the chat AI model and generates a response message.
[0627] Step 11:
[0628] Chat AI model: Analyzes messages and generates appropriate response messages.
[0629] Step 12:
[0630] Server: Input the response message into the speech generation model and generate speech data.
[0631] Step 13:
[0632] Speech generation model: Analyzes text data and generates speech data.
[0633] Step 14:
[0634] Server: Sends the generated audio data to the user's device.
[0635] Step 15:
[0636] Terminal: Plays the audio data and the user hears the character's response.
[0637] Step 16:
[0638] User: Set preferences to customize the tone of the dialogue and character reactions.
[0639] Step 17:
[0640] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0641] This series of steps allows users to enjoy natural interactions with their chosen fictional characters.
[0642] Example 1
[0643] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0644] With conventional dialogue systems, it has been difficult to achieve natural and personalized dialogue with fictional characters. It has also been difficult to properly reflect the user's ratings and preferences and respond individually. This has resulted in inconsistent dialogue experiences with the selected characters, leading to low satisfaction. Furthermore, the process of voice message recognition and response generation is complex, making system setup and use time-consuming.
[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0646] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for sending the generated voice data to the user, means for comparing authentication information entered by the user with a database and displaying an access portal if authentication is successful, means for loading data of a virtual character selected by the user and performing initial settings, means for speech recognition to convert voice messages into text, and means for saving conversation history and user settings in a database and loading them as needed, thereby enabling natural and personalized interactions with fictional characters.
[0647] The "means for receiving an input message from a user" refers to a device or program that has the function of transmitting a text message or voice message input by a user to a server.
[0648] "Means for initializing a chat AI model using a natural language processing algorithm and generating a response message" refers to a device or program that has the function of initiating the operation of a chat AI model using natural language processing technology and generating an appropriate response to a message from a user.
[0649] The "means for initializing a speech generation model and generating speech data" refers to a device or program having the function of operating a speech generation model for converting a text message into speech data and converting the generated text response into speech data.
[0650] The "means for transmitting generated voice data to a user" refers to a device or program having the function of transmitting voice data generated by the server to a user's terminal.
[0651] "Means for checking the authentication information entered by the user against a database and displaying an access portal if authentication is successful" refers to a device or program that has the function of checking the authentication information entered by the user (such as a user name and password) against a database to perform authentication, and displaying the main screen if authentication is successful.
[0652] "Means for loading the data of a virtual character selected by the user and performing initial settings" refers to a device or program that has the function of loading the data of a character selected by the user and setting the corresponding chat AI model and voice generation model.
[0653] The "voice recognition means for converting a voice message into text" is a device or program that has the function of converting a voice message uttered by a user into text format.
[0654] "Means for saving conversation history and user settings in a database and loading them as needed" refers to a device or program that has the function of recording the conversation history and user settings in a database and referencing this information during subsequent conversations.
[0655] This invention relates to a system that allows users to interact with fictional characters. The system consists of a server and a user terminal, and utilizes a chat AI model and a voice generation model to realize natural and personalized interactions.
[0656] System configuration
[0657] The system mainly consists of the following components:
[0658] 1. Server: A high-performance computer is used to process the chat AI model and voice generation model.
[0659] 2. Terminal: The device through which the user accesses the service, including smartphones and PCs.
[0660] 3. Database: Stores authentication information, character settings, conversation history, and user preferences.
[0661] Hardware and software used
[0662] Hardware: High-performance servers (e.g., cloud service provider servers), user devices (smartphones, PCs)
[0663] Software: Database (e.g., MySQL), chat AI model (e.g., GPT-3), speech generation model (e.g., Google Cloud Text-to-Speech), speech recognition method (e.g., Google Cloud Speech-to-Text)
[0664] Explanation of program processing
[0665] User authentication
[0666] The user accesses the system's login screen from a terminal and enters their username and password. The terminal encrypts this authentication information and sends it to the server, which checks it against a database and, if authentication is successful, displays the access portal. This process uses a MySQL database.
[0667] Character Selection
[0668] The user selects the character they want to interact with from the access portal. The device sends the character information to the server, which then loads the chat AI model (e.g., GPT-3) and voice generation model (e.g., Google Cloud Text-to-Speech) corresponding to the selected character and performs initial setup.
[0669] Conversational Interactions
[0670] Users send text or voice messages from their devices. When a voice message is sent, the device converts the speech to text using Google Cloud Speech-to-Text. The device then sends the text message to the server. The server inputs the message into a chat AI model and generates a reply message. The generated reply message is then input into a voice generation model, which generates voice data. The server then sends this voice data to the user's device, which then plays it back.
[0671] Customization features
[0672] Users can customize the tone of the conversation and the character's reactions in the settings screen, and this information is stored in a database on the server and applied to future interactions.
[0673] Specific examples
[0674] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[0675] User: Texts "How was your day?"
[0676] Terminal: Sends a text message to the server.
[0677] Server: Passes the message to the chat AI model (GPT-3) and generates a response message.
[0678] Chat AI model: Generates the response "Today was a really fun day."
[0679] Speech generation model: Converts the response message into speech data.
[0680] Server: Sends the audio data to the user's device.
[0681] Terminal: Plays the audio data and the user hears the character's response.
[0682] Prompt Sentence Examples
[0683] "Ask Character A, 'How was your day?'"
[0684] The invention allows users to enjoy natural and personalized interactions with fictional characters.
[0685] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0686] Step 1:
[0687] User: Access the device's login screen and enter your username and password.
[0688] Input: Username and Password.
[0689] Output: The authentication information is sent to the device.
[0690] Step 2:
[0691] On the device: The authentication information entered by the user is encrypted and sent to the server.
[0692] Input: Encrypted credentials.
[0693] Output: Authentication information sent to the server.
[0694] Step 3:
[0695] Server: Compares the user information stored in the database, and if authentication is successful, displays the access portal; if authentication fails, generates and returns an error message.
[0696] Input: Encrypted credentials.
[0697] Output: The authentication result (success or failure).
[0698] Step 4:
[0699] User: Select the desired character from the character list displayed on the dashboard.
[0700] Input: The ID of the character selected by the user.
[0701] Output: Character selection information is sent to the device.
[0702] Step 5:
[0703] Device: Sends the selected character information to the server.
[0704] Input: Character ID.
[0705] Output: Character information is sent to the server.
[0706] Step 6:
[0707] Server: Loads the chat AI model and voice generation model corresponding to the selected character and performs initial settings.
[0708] Input: Character ID.
[0709] Output: The chat AI model and speech generation model are initialized.
[0710] Step 7:
[0711] User: Send a text or voice message to the character saying, "How was your day?"
[0712] Input: The user's text or voice message.
[0713] Output: The message is sent to the terminal.
[0714] Step 8:
[0715] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0716] Input: The user's voice message.
[0717] Output: A text message.
[0718] Step 9:
[0719] Terminal: Sends a text message to the server.
[0720] Input: A text message.
[0721] Output: A message is sent to the server.
[0722] Step 10:
[0723] Server: Inputs the received text message into the chat AI model and generates a response message.
[0724] Input: A text message.
[0725] Output: The response message.
[0726] Step 11:
[0727] Chat AI model: Analyzes messages and generates a response message such as "I had a great day today."
[0728] Input: The user's text message.
[0729] Output: The response message.
[0730] Step 12:
[0731] Server: Input the response message into the speech generation model and generate speech data.
[0732] Input: The response message.
[0733] Output: Audio data.
[0734] Step 13:
[0735] Speech generation model: Receives the response message and generates speech data.
[0736] Input: The response message.
[0737] Output: Audio data.
[0738] Step 14:
[0739] Server: Sends the generated audio data to the user's device.
[0740] Input: Audio data.
[0741] Output: The audio data is sent to the device.
[0742] Step 15:
[0743] Terminal: Plays the audio data and the user hears the character's response.
[0744] Input: Audio data.
[0745] Output: Playback of audio data.
[0746] (Application example 1)
[0747] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0748] In virtual environments, users are expected to have a natural experience through interactions with fictional characters. However, conventional systems have made it difficult for users to interact with characters in real time within a virtual environment and receive product explanations and guidance, making it difficult to provide a truly immersive experience. The present invention aims to solve these problems by providing a system that allows users to interact with fictional characters in a virtual environment in a natural and interactive manner, and to receive product explanations and guidance.
[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0750] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for transmitting the generated voice data to the user, and means for reproducing a fictional character in the virtual environment based on a selection by the user and providing guidance or product explanations, thereby enabling the user to have natural conversations with the fictional character in the virtual environment in real time and receive product explanations or guidance.
[0751] "Means for receiving input messages from a user" means a mechanism by which the system receives messages input by a user when the user interacts with a fictional character within a virtual environment.
[0752] A "chat AI model" is an artificial intelligence model built on a natural language processing algorithm used to generate natural response messages in response to user input messages.
[0753] A "voice generation model" is a model used to convert a text response message generated by a chat AI model into voice data.
[0754] A "virtual environment" is a computer-generated virtual space experienced by a user, for example, using a VR headset or smart device.
[0755] A "fictional character" is a fictional character that appears in fictional works such as movies, anime, and games, and whose role is to provide an experience through interactions with users.
[0756] "User authentication" is the process of identifying and verifying a user's identity when accessing a system.
[0757] The "dashboard" is the interface that users can access after authentication, and is the main screen for selecting fictional characters.
[0758] "Real-time response" is a function that generates a response immediately to user input, allowing the character to respond immediately.
[0759] A "database" is a system that stores information such as user settings and conversation history and can read that information as needed.
[0760] "Guide and product information" refers to an explanation or introduction given to a user by a fictional character, providing product details and purchasing advice.
[0761] A specific embodiment of the present invention is to build a system that allows users to naturally interact with fictional characters in a virtual environment. This system is composed of a user terminal and a cloud server.
[0762] System Configuration
[0763] Hardware and Software
[0764] User devices: smartphones (iOS or Android), VR headsets (e.g., Oculus Quest 2), smart glasses.
[0765] Server: Cloud service (e.g. AWS EC2).
[0766] Database: A cloud database service (e.g. AWS RDS).
[0767] AI models: Chat AI models (e.g., OpenAI GPT-4), speech generation models (e.g., Google Wavenet).
[0768] Program processing
[0769] 1. User authentication:
[0770] When a user logs in to the app, the device sends authentication information to the server.
[0771] The server checks the authentication information against the database, and if authentication is successful, displays the dashboard on the terminal.
[0772] 2. Character Selection:
[0773] Users select a fictional character from the dashboard.
[0774] The device sends the selected character information to the server, and the server loads the chat AI model and voice generation model for the character and performs initial settings.
[0775] 3. Conversational Interactions:
[0776] When the user types "Hello, what do you recommend today?", the device sends this message to the server.
[0777] The server passes the input message to the chat AI model to generate a response message.
[0778] The chat AI model generates a response message saying, "Hello! This month's recommendation is new earphones," and returns it to the server.
[0779] The server sends a response message to the speech generation model to generate speech data.
[0780] The generated audio data is returned to the terminal, which then plays the audio.
[0781] 4. Enhanced user experience:
[0782] Users can enjoy a virtual experience in which fictional characters act as guides, showing them around the store and explaining products.
[0783] Conversation history and user preferences are stored in a database for later reuse.
[0784] Specific examples
[0785] For example, if the user selects "Fictional Character B":
[0786] 1. Put on the VR headset.
[0787] 2. A user sends a text message asking, "What are your recommendations in the store?"
[0788] 3. The chat AI model responds, "The recommended product is the newly released smartwatch."
[0789] 4. The speech generation model converts the response into speech.
[0790] 5. Playback on the user's device.
[0791] In this way, it is possible to provide users with an engaging experience by realizing natural interactions and real-time responses with fictional characters within a virtual environment.
[0792] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0793] Step 1:
[0794] User authentication
[0795] The user logs in to the app and enters their user ID and password.
[0796] The terminal transmits the input authentication information to the server.
[0797] The server checks the authentication information against the database, and if authentication is successful, generates dashboard information along with an authentication success message and sends it to the terminal. If authentication fails, it returns an error message.
[0798] The device displays the received dashboard information to the user and waits for the next action.
[0799] Step 2:
[0800] Character Selection
[0801] Users select the fictional character they want from the dashboard.
[0802] The terminal transmits the selected character information to the server and displays a confirmation message.
[0803] The server loads the chat AI model and voice generation model for the character, performs initial setup, and returns a message to the device indicating that the setup is complete.
[0804] The terminal will display an initialization complete message to the user and wait for further action.
[0805] Step 3:
[0806] Interaction Start
[0807] The user types a text message (e.g., "Hello, what do you recommend today?") or a voice message to the selected character.
[0808] If the message is a voice message, the terminal performs voice recognition processing to convert it into text, and then sends the text message to the server.
[0809] The server inputs the received text message into the chat AI model and generates a response message.
[0810] The chat AI model generates a response message such as "Hello! This month's recommendation is new earphones" and returns it to the server.
[0811] Step 4:
[0812] Generating a response message
[0813] The server sends the response message received from the chat AI model to the voice generation model, which generates voice data.
[0814] The speech generation model analyzes the text message and generates corresponding speech data, which is returned to the server.
[0815] Step 5:
[0816] Sending and playing reply messages
[0817] The server transmits the generated voice data to the terminal.
[0818] The device plays back the received audio data, allowing the user to hear the fictional character's response.
[0819] Along with the response message, it displays UI elements to determine the interaction status and next action.
[0820] Step 6:
[0821] Save conversation history and settings
[0822] The server stores conversation history and user settings (e.g. character customization) in a database.
[0823] The database stores the necessary information to allow users to access their conversation history and preferences the next time they log in.
[0824] This series of processes allows users to have natural conversations with fictional characters in a virtual environment, receive responses in real time, and enjoy product guides and explanations in a virtual store.
[0825] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0826] This system recreates fictional characters and allows natural dialogue with users. It also incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly, providing more personalized interactions. The system consists of a server and a user terminal.
[0827] System Overview
[0828] The system's server handles the main processing and responds to user requests using a chat AI model, a voice generation model, and an emotion engine. Users can access the system via their devices and enjoy interacting with the characters.
[0829] User authentication and character selection
[0830] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[0831] 2. Terminal: Sends the entered authentication information to the server.
[0832] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[0833] 4. User: Select the fictional character of your choice from the dashboard.
[0834] 5. Terminal: Sends the selected character information to the server.
[0835] 6. Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[0836] Conversational Interactions
[0837] 1. User: Send a text or voice message to the character from your device.
[0838] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[0839] 3. Device: Sends a text message to the server.
[0840] 4. Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[0841] 5. Emotion Engine: Recognizes user emotions from text messages and generates emotional status.
[0842] 6. Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[0843] 7. Chat AI model: Generates a response message and returns it to the server.
[0844] 8. Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[0845] 9. Speech generation model: Generates speech data based on text data and emotional status.
[0846] 10. Server: Sends the generated voice data to the user's device.
[0847] 11. Terminal: Plays the audio data and the user hears the character's response.
[0848] Specific examples
[0849] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[0850] 1. User: Texts "I'm so tired today."
[0851] 2. Server: Receives the message and passes it to the emotion engine.
[0852] 3. Emotion engine: Analyzes the message and generates an emotional status that identifies the user as tired.
[0853] 4. Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been tough, take a break."
[0854] 5. Speech generation model: Receives a response message and generates speech data in a gentle tone according to the emotional status.
[0855] 6. Server: Sends the audio data to the user's device.
[0856] 7. Terminal: Play the audio data and the character responds in a gentle tone, saying, "That must have been hard. Take a break."
[0857] Customization features
[0858] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[0859] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0860] This system allows users to enjoy natural and emotional interactions with their chosen fictional characters, and individual settings can provide a more personalized experience, allowing for deeper emotional fulfillment for people known as fictiosexuals.
[0861] The processing flow will be explained below.
[0862] This invention is a system that recreates fictional characters and enables natural dialogue with users, and by incorporating an emotion engine that recognizes the user's emotions and adjusts the response content, it provides more personalized interactions. The specific processing flow of the system is explained below in steps.
[0863] User authentication and character selection
[0864] Step 1:
[0865] User: Accesses the system login screen from a terminal and enters a username and password.
[0866] Step 2:
[0867] Terminal: Sends the entered authentication information to the server.
[0868] Step 3:
[0869] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[0870] Step 4:
[0871] Users: Select the fictional character of their choice from the dashboard.
[0872] Step 5:
[0873] Device: Sends the selected character information to the server.
[0874] Step 6:
[0875] Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[0876] Conversational Interactions
[0877] Step 7:
[0878] User: Send a text or voice message from your device to the character.
[0879] Step 8:
[0880] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[0881] Step 9:
[0882] Terminal: Sends a text message to the server.
[0883] Step 10:
[0884] Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[0885] Step 11:
[0886] Emotion Engine: Recognizes user emotions from text messages and generates emotional statuses.
[0887] Step 12:
[0888] Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[0889] Step 13:
[0890] Chat AI model: Analyzes messages and generates appropriate response messages.
[0891] Step 14:
[0892] Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[0893] Step 15:
[0894] Speech generation model: Generates speech data based on text data and emotional status.
[0895] Step 16:
[0896] Server: Sends the generated audio data to the user's device.
[0897] Step 17:
[0898] Terminal: Plays the audio data and the user hears the character's response.
[0899] Specific examples
[0900] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[0901] Step 1:
[0902] User: Texts "I'm so tired today."
[0903] Step 2:
[0904] Server: Receives messages and passes them to the emotion engine.
[0905] Step 3:
[0906] Emotion engine: Analyzes messages and generates an emotional status that identifies the user as tired.
[0907] Step 4:
[0908] Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been hard, take a break."
[0909] Step 5:
[0910] Speech generation model: Receives a response message and generates voice data in a gentle tone according to the emotional status.
[0911] Step 6:
[0912] Server: Sends the audio data to the user's device.
[0913] Step 7:
[0914] Device: Plays the audio data and the character responds in a gentle tone, "That must have been hard, take a break."
[0915] Customization features
[0916] Step 1:
[0917] User: You can set preferences to customize the tone of the dialogue and character reactions.
[0918] Step 2:
[0919] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[0920] This allows the system to realize natural and emotional interactions with fictional characters selected by the user, making it possible to more deeply satisfy the emotions of people known as fictiosexuals.
[0921] Example 2
[0922] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0923] Conventional dialogue systems have difficulty in engaging in natural conversations with fictional characters, limiting their ability to provide personalized responses based on the user's emotions. Furthermore, the accuracy of speech recognition and speech synthesis is low, making it difficult to generate speech that takes emotions into account. Furthermore, systems lack a mechanism for centrally managing user authentication information and settings and saving dialogue history to provide a better user experience. Therefore, a new dialogue system is needed to improve user satisfaction.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0925] In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue generation model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice synthesis model to convert the response message into voice data and generating voice data, means for utilizing an emotion analysis engine to analyze the user's emotion from the input message, means for generating a response message taking into account the emotional status obtained by the emotion analysis engine, means for utilizing the voice synthesis model to generate voice data with a tone corresponding to the emotional status, and means for transmitting the generated voice data to the user. This makes it possible to generate a response message and voice data that take into account the user's emotion and provide a natural and personalized dialogue with a fictional character.
[0926] "User" means any individual or entity that accesses the System and interacts with fictional characters.
[0927] "Input Message" refers to text or voice information sent by a user to a system.
[0928] "Natural language processing algorithms" refer to technologies for understanding a user's input message and generating an appropriate response.
[0929] "Dialogue generation model" refers to a machine learning model that uses a natural language processing algorithm to generate a response message to a user's input message.
[0930] "Response message" refers to text information that the dialogue generation model generates in response to a user's input message.
[0931] "Speech synthesis model" refers to the technology for converting text response messages into speech data.
[0932] "Audio Data" refers to an audio file or audio stream generated by a speech synthesis model.
[0933] "Sentiment analysis engine" refers to technology for analyzing emotions from a user's input message and generating an emotional status.
[0934] "Emotional status" refers to information that expresses a user's emotions numerically or categorically, as determined by the emotion analysis engine.
[0935] "Authentication Information" refers to information such as username and password that a user uses to log into a system.
[0936] "Dashboard" refers to the interface that allows users to select characters and change settings.
[0937] "Fictional Character" refers to any imaginary or fictional person or entity that is recreated within the system.
[0938] "Voice recognition means" refers to technology used to convert a voice message sent by a user into a text message.
[0939] "Conversation history" refers to a record of the dialogue between a user and a character.
[0940] "Database" refers to a management system for systematically storing information necessary for system operation.
[0941] This invention is a system that allows users to have natural and emotional conversations with fictional characters. This system analyzes the user's emotions and appropriately adjusts the content and tone of the conversation to provide a more personalized conversation. Specifically, the system uses the following hardware and software to process data and perform calculations.
[0942] Hardware and software used
[0943] server
[0944] Database Management System (DBMS): Stores user authentication information and configuration data, typically using MySQL or PostgreSQL.
[0945] Dialogue generation model: Generates dialogue. As an example, GPT-3, a machine learning model, is used.
[0946] Speech synthesis model: Converts the response message into audio data. For example, speech synthesis APIs such as Google Cloud Text-to-Speech and IBM Watson Text to Speech are used.
[0947] Sentiment analysis engine: Recognizes user emotions. For example, Microsoft Azure's Sentiment Analysis is used.
[0948] Terminal
[0949] User Interface (UI): A screen where users can interact with fictional characters. Most commonly, it is a web app that is made up of HTML / CSS and JavaScript.
[0950] Speech recognition method: Converts voice input into text. Google Cloud Speech-to-Text and IBM Watson Speech to Text are used.
[0951] User authentication and character selection
[0952] The system begins when a user logs in using their authentication information. The user accesses the system's login screen from their terminal and enters their username and password. The input information is sent from the terminal to the server, which then checks it against a database to authenticate the user.
[0953] After successful authentication, the user is redirected to a dashboard where they can select their desired fictional character. The selected character information is sent from the device to the server, which then loads and initializes the character's corresponding dialogue generation model, speech synthesis model, and sentiment analysis engine.
[0954] Conversational Interactions
[0955] When a user starts a dialogue with a character, an input message is sent from the terminal to the server. If this input message is a voice message, the terminal converts it into text using a voice recognition means.
[0956] The received text message is processed by the server and passed to the sentiment analysis engine, which analyzes the message and generates the user's emotional status. The emotional status is input into the dialogue generation model, and a response message is generated that takes the user's emotions into account.
[0957] The generated response message and emotional status are passed to a speech synthesis model to generate speech data, including tones, which is then sent from the server to the user's device and played back on the device.
[0958] Customization features
[0959] The system provides users with settings to customize the tone of dialogue and character reactions. Users can adjust various settings by manipulating sliders and checkboxes on the settings page. These settings are saved in a database by the server and applied the next time a dialogue is held.
[0960] Specific examples
[0961] For example, if a user sends a message to "Fictional Character A" saying "I'm really tired today," the system operates as follows: First, the user sends a text message saying "I'm really tired today." The server receives this message and passes it to the emotion analysis engine. The emotion analysis engine recognizes that the user is tired and generates an emotion status.
[0962] Next, the server inputs the emotional status into a dialogue generation model and generates a response message that matches the user's emotion, "That must have been hard. Take a break." The speech synthesis model receives this response message and generates voice data in a gentle tone according to the emotional status. Finally, the server sends the voice data to the user's device, which plays it back, with the character responding in a gentle tone, "That must have been hard. Take a break."
[0963] Prompt Sentence Examples
[0964] If a user sends a message to fictional character A saying, "I'm really tired today," please explain in detail how the emotion engine generates the emotion status and what response the dialogue generation model makes.
[0965] As described above, the present invention is a system that allows users to enjoy natural and emotional interactions with fictional characters of their choice, providing a more personalized experience through individual settings.
[0966] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0967] Divide the program's processing flow into processing steps
[0968] Step 1:
[0969] Step 2:
[0970] Step 3:
[0971] ...
[0972] Specific explanations for each step
[0973] Step 1: User authentication and login
[0974] Input: Username, Password
[0975] Output: Authentication result (success or failure), dashboard screen or error message
[0976] Operation:
[0977] 1. User: Enter your username and password on the device login screen.
[0978] 2. Terminal: Sends the input information to the server as a POST request.
[0979] 3. Server: Query the database to verify the authentication information. If authentication is successful, return the dashboard screen; if authentication is unsuccessful, return an error message.
[0980] Step 2: Choose your character
[0981] Input: Character ID
[0982] Output: Character data, initialization result
[0983] Operation:
[0984] 1. User: Select the fictional character of your choice from the dashboard.
[0985] 2. Device: Send the selected character ID to the server as an AJAX request.
[0986] 3. Server: Loads the dialogue generation model, speech synthesis model, and emotion analysis engine corresponding to the character and performs initial configuration.
[0987] Step 3: Send the message
[0988] Input: Text or voice message
[0989] Output: Text message (in case of voice)
[0990] Operation:
[0991] 1. User: Send a text or voice message from your device to the character.
[0992] 2. Device: When a voice message is sent, a speech recognition tool is used to convert the voice into text, the audio file is sent to a cloud API, and the text is returned.
[0993] Step 4: Sentiment Analysis
[0994] Input: Text message
[0995] Output: Emotional status
[0996] Operation:
[0997] 1. Terminal: Sends the converted text message to the server.
[0998] 2. Server: Inputs the received text message into the sentiment analysis engine to analyze the user's sentiment. The sentiment engine analyzes the message and generates a sentiment score.
[0999] Step 5: Generate a response message
[1000] Input: Text message, emotional status
[1001] Output: Response message
[1002] Operation:
[1003] 1. Server: Inputs the emotional status into the dialogue generation model and generates a response message that takes the user's emotions into account. The generative AI model analyzes the text message and emotional status and generates an appropriate response.
[1004] Step 6: Generate audio data
[1005] Input: Response message, emotional status
[1006] Output: Audio data
[1007] Operation:
[1008] 1. Server: Input the response message and emotional status into the speech synthesis model and generate speech data with a tone corresponding to the emotion. The speech synthesis API generates an audio file with a tone based on the emotion.
[1009] Step 7: Playing back audio data
[1010] Input: Audio data
[1011] Output: Play audio
[1012] Operation:
[1013] 1. Server: Sends the generated voice data to the user's device.
[1014] 2. Device: The audio data is played and the user hears the character's response. The device's audio player plays the audio data.
[1015] This is the flow of the program's processing. The specific operations, inputs, and outputs at each step are explained in detail, and it contains all the elements necessary for users to enjoy natural and emotional interactions with fictional characters.
[1016] (Application example 2)
[1017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1018] Conventional chat systems and voice dialogue systems have difficulty responding to user emotions during interactions, resulting in a limited user experience. In particular, when dealing with customers in physical stores, flexible responses based on customer emotions are required, but the means to achieve this are not fully developed. As a result, there is a risk of customer satisfaction decreasing.
[1019] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue AI model using a natural language processing algorithm for responding to the input message and generating a response message, means for initializing a voice generation algorithm for converting the response message into voice data and generating voice data, means for adding emotion to the response message using an emotion analysis engine that analyzes the user's emotion, and means for transmitting the generated voice data to the user. This enables natural and personalized dialogue that takes the user's emotion into consideration.
[1020] "Means for receiving input messages from a user" means a system or function for obtaining text or voice messages sent by a user from a terminal.
[1021] A "dialogue AI model using natural language processing algorithms" is an algorithm or model that analyzes a user's input message and generates a response accordingly.
[1022] "Speech generation algorithm" refers to an algorithm or technique for converting a text response message into voice data.
[1023] An "emotion analysis engine" is an analysis system that identifies emotions from the user's input message and reflects those emotions in the response message.
[1024] "Means for verifying user authentication information" refers to a system or function that checks authentication information such as a user's ID and password to determine whether the user is legitimate.
[1025] The "means for displaying a dashboard" refers to a system or function for providing an operation screen that is displayed after a user logs in to the system.
[1026] A "means for loading fictional character data" is a system or function that loads data or settings related to a specific fictional character into memory and makes them available for use.
[1027] "Means for analyzing the user's emotional status and reflecting it in the initial settings" refers to a system or function for analyzing the user's emotions and reflecting the results in the character's settings and behavior.
[1028] "Speech recognition means" refers to technology or systems that convert the user's speech into text.
[1029] "Means for storing conversation history and user settings in a database" refers to a system or function that records the content of conversations with users and user settings (such as customization information) in a database, allowing them to be searched or referenced when necessary.
[1030] A "means for customizing fictional character responses" is a system or feature that allows a fictional character's words and actions to be flexibly changed based on the user's settings and choices.
[1031] In this invention, a natural, emotion-conscious dialogue between a user and a fictional character can be realized using the following system configuration and processing procedure.
[1032] Overall system configuration
[1033] The server contains the following main modules:
[1034] 1. Input message receiving module
[1035] Receive text and voice messages from user devices.
[1036] 2. Dialogue AI Model
[1037] It uses natural language processing algorithms to respond to input messages from users.
[1038] 3. Speech Generation Algorithm
[1039] The response message is converted into voice data.
[1040] 4. Sentiment Analysis Engine
[1041] Analyze the user's emotions from the input message and reflect them in the response message.
[1042] 5. User Authentication Module
[1043] Verify the user's credentials and display the dashboard if authentication is successful.
[1044] 6. Character Data Load Module
[1045] Loads a fictional character of the user's choice and performs initial setup.
[1046] 7. Speech Recognition Module
[1047] Convert voice messages to text.
[1048] 8. Database Management Module
[1049] Conversation history and user preferences are stored in a database and loaded as needed.
[1050] 9. Character Reaction Customization Module
[1051] Customize character reactions based on user choices and settings.
[1052] Hardware and Software
[1053] The hardware and software used are as follows:
[1054] Server hardware: A server with a powerful processor and sufficient memory and storage. For example, you can use cloud services such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[1055] Software: Generative AI models such as GPT-3 are used for natural language processing algorithms. WaveNet and Tacotron are used for speech generation algorithms. Sentiment analysis libraries (Python text2emotion, transformers, etc.) are used for sentiment analysis.
[1056] Specific explanation of the process
[1057] server
[1058] 1. The input message receiving module receives a message from the user, which can be text or voice.
[1059] 2. The speech recognition module converts the voice message into text.
[1060] 3. The converted text is input into a sentiment analysis engine to analyze the user's sentiment.
[1061] 4. The analyzed emotional status is input into a dialogue AI model to generate a response message that takes emotions into account.
[1062] 5. The response message is converted into voice data by a voice generation algorithm.
[1063] 6. The generated voice data is sent to the user terminal.
[1064] 7. The database management module stores conversation history and user settings for future use.
[1065] 8. Character reaction customization module modifies character reactions based on user settings.
[1066] Examples and prompts
[1067] For example, if a customer asks "What products do you recommend?" in a physical store, the specific processing is as follows:
[1068] User: Ask "What products do you recommend?"
[1069] Server: The sentiment analysis engine analyzes the sentiment of the question "What product do you recommend?", and the dialogue AI model generates a response message.
[1070] The voice generation algorithm responds, "Today's recommendation is fresh apples," and the generated voice data is sent to the user's device.
[1071] Example prompt sentence:
[1072] "Please enter a user message: What products do you recommend?"
[1073] "Please perform sentiment analysis on this message and generate an appropriate response."
[1074] In this way, a natural and emotional interaction with the user can be achieved.
[1075] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1076] Step 1:
[1077] The user inputs a voice message or text message from a device such as a smartphone or tablet, and the input message is sent from the device to the server.
[1078] Step 2:
[1079] The server invokes a speech recognition module to analyze the received message. When a voice message is sent, the speech recognition module converts the voice data into text data. The text data is used as input and the speech recognition result is obtained as output.
[1080] Step 3:
[1081] The server inputs the text message into the sentiment analysis engine. The sentiment analysis engine analyzes the emotional status from the user's input message. In this step, the sentiment analysis engine analyzes the text data and extracts the emotional status from the user's text message. The input is the text data and the output is the emotional status.
[1082] Step 4:
[1083] The server inputs the analyzed emotional status into a dialogue AI model to generate a response message that takes the user's emotions into account. The dialogue AI model is based on a generative AI model, which receives the emotional status and text message as input and generates a response message as output. This process results in an appropriate response that corresponds to the user's input and emotions.
[1084] Step 5:
[1085] The server sends the generated response message to a voice generation algorithm to generate voice data. The voice generation algorithm receives the response message and the emotional status as input and outputs the voice data in a tone that reflects the emotion.
[1086] Step 6:
[1087] The server transmits the generated voice data to the user's device, which then plays the received voice data and provides the character's response to the user, allowing the user to experience voice responses that reflect their emotions.
[1088] Step 7:
[1089] The server stores the conversation history and user settings in a database. The database management module uses the conversation data and settings information as input and obtains the saved data as output. This data can be used as a reference for future conversations.
[1090] Step 8:
[1091] The user customizes the character's reactions as needed. The server's character reaction customization module receives the user's settings and generates the character's reactions based on them. In this step, the user's input settings are processed and the customized reactions are output.
[1092] The above are the processing steps of the system that realizes this invention, which allows the user to have a natural conversation experience that takes emotions into consideration.
[1093] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1094] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1095] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1096] [Third embodiment]
[1097] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1098] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1099] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1100] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1101] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1102] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1103] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1104] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1105] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1106] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1107] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1108] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1109] The present invention is a system for recreating fictional characters, and is designed to allow users to have natural conversations with the characters. Specific embodiments for implementing this system are as follows.
[1110] System Overview
[1111] This system consists of a server and a user terminal. Users can access the system through their terminal and enjoy interacting with characters. The server is responsible for the main processing and responds to user requests using a chat AI model and a voice generation model.
[1112] User authentication and character selection
[1113] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[1114] 2. Terminal: Sends the entered authentication information to the server.
[1115] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[1116] 4. User: Select the fictional character of your choice from the dashboard.
[1117] 5. Terminal: Sends the selected character information to the server.
[1118] 6. Server: Loads the chat AI model and voice generation model corresponding to the character and performs initial settings.
[1119] Conversational Interactions
[1120] 1. User: Send a text or voice message to the character from your device.
[1121] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[1122] 3. Device: Sends a text message to the server.
[1123] 4. Server: Inputs the received text message into the chat AI model and generates a response message.
[1124] 5. Chat AI model: Generates a response message and returns it to the server.
[1125] 6. Server: Input the response message into the speech generation model and generate speech data.
[1126] 7. Speech generation model: Generates speech data and returns it to the server.
[1127] 8. Server: Sends the generated voice data to the user's device.
[1128] 9. Terminal: Plays the audio data and the user hears the character's response.
[1129] Specific examples
[1130] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[1131] 1. User: Sends a text message saying, "How was your day?"
[1132] 2. Server: Receives messages and passes them to the chat AI model.
[1133] 3. Chat AI model: Analyzes the message and generates a response message such as, "Today was a really fun day."
[1134] 4. Speech generation model: Receives the response message and generates speech data.
[1135] 5. Server: Sends the audio data to the user's device.
[1136] 6. Terminal: Plays the audio data and the user hears the character's response.
[1137] Customization features
[1138] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[1139] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1140] The system allows users to enjoy natural dialogue with the characters they select, and can provide a more personalized experience through individual settings, which can satisfy the feelings of people known as fictiosexuals.
[1141] The processing flow will be explained below.
[1142] Step 1:
[1143] User: Accesses the system login screen and enters a username and password.
[1144] Step 2:
[1145] Terminal: Sends the entered authentication information to the server.
[1146] Step 3:
[1147] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[1148] Step 4:
[1149] Users: Select the fictional character of their choice from the dashboard.
[1150] Step 5:
[1151] Device: Sends the selected character information to the server.
[1152] Step 6:
[1153] Server: Loads the chat AI model and voice generation model corresponding to the characters and performs initial settings.
[1154] Step 7:
[1155] User: Send a text or voice message from your device to the character.
[1156] Step 8:
[1157] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1158] Step 9:
[1159] Terminal: Sends a text message to the server.
[1160] Step 10:
[1161] Server: Inputs the received text message into the chat AI model and generates a response message.
[1162] Step 11:
[1163] Chat AI model: Analyzes messages and generates appropriate response messages.
[1164] Step 12:
[1165] Server: Input the response message into the speech generation model and generate speech data.
[1166] Step 13:
[1167] Speech generation model: Analyzes text data and generates speech data.
[1168] Step 14:
[1169] Server: Sends the generated audio data to the user's device.
[1170] Step 15:
[1171] Terminal: Plays the audio data and the user hears the character's response.
[1172] Step 16:
[1173] User: Set preferences to customize the tone of the dialogue and character reactions.
[1174] Step 17:
[1175] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1176] This series of steps allows users to enjoy natural interactions with their chosen fictional characters.
[1177] Example 1
[1178] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1179] With conventional dialogue systems, it has been difficult to achieve natural and personalized dialogue with fictional characters. It has also been difficult to properly reflect the user's ratings and preferences and respond individually. This has resulted in inconsistent dialogue experiences with the selected characters, leading to low satisfaction. Furthermore, the process of voice message recognition and response generation is complex, making system setup and use time-consuming.
[1180] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1181] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for sending the generated voice data to the user, means for comparing authentication information entered by the user with a database and displaying an access portal if authentication is successful, means for loading data of a virtual character selected by the user and performing initial settings, means for speech recognition to convert voice messages into text, and means for saving conversation history and user settings in a database and loading them as needed, thereby enabling natural and personalized interactions with fictional characters.
[1182] The "means for receiving an input message from a user" refers to a device or program that has the function of transmitting a text message or voice message input by a user to a server.
[1183] "Means for initializing a chat AI model using a natural language processing algorithm and generating a response message" refers to a device or program that has the function of initiating the operation of a chat AI model using natural language processing technology and generating an appropriate response to a message from a user.
[1184] The "means for initializing a speech generation model and generating speech data" refers to a device or program having the function of operating a speech generation model for converting a text message into speech data and converting the generated text response into speech data.
[1185] The "means for transmitting generated voice data to a user" refers to a device or program having the function of transmitting voice data generated by the server to a user's terminal.
[1186] "Means for checking the authentication information entered by the user against a database and displaying an access portal if authentication is successful" refers to a device or program that has the function of checking the authentication information entered by the user (such as a user name and password) against a database to perform authentication, and displaying the main screen if authentication is successful.
[1187] "Means for loading the data of a virtual character selected by the user and performing initial settings" refers to a device or program that has the function of loading the data of a character selected by the user and setting the corresponding chat AI model and voice generation model.
[1188] The "voice recognition means for converting a voice message into text" is a device or program that has the function of converting a voice message uttered by a user into text format.
[1189] "Means for saving conversation history and user settings in a database and loading them as needed" refers to a device or program that has the function of recording the conversation history and user settings in a database and referencing this information during subsequent conversations.
[1190] This invention relates to a system that allows users to interact with fictional characters. The system consists of a server and a user terminal, and utilizes a chat AI model and a voice generation model to realize natural and personalized interactions.
[1191] System configuration
[1192] The system mainly consists of the following components:
[1193] 1. Server: A high-performance computer is used to process the chat AI model and voice generation model.
[1194] 2. Terminal: The device through which the user accesses the service, including smartphones and PCs.
[1195] 3. Database: Stores authentication information, character settings, conversation history, and user preferences.
[1196] Hardware and software used
[1197] Hardware: High-performance servers (e.g., cloud service provider servers), user devices (smartphones, PCs)
[1198] Software: Database (e.g., MySQL), chat AI model (e.g., GPT-3), speech generation model (e.g., Google Cloud Text-to-Speech), speech recognition method (e.g., Google Cloud Speech-to-Text)
[1199] Explanation of program processing
[1200] User authentication
[1201] The user accesses the system's login screen from a terminal and enters their username and password. The terminal encrypts this authentication information and sends it to the server, which checks it against a database and, if authentication is successful, displays the access portal. This process uses a MySQL database.
[1202] Character Selection
[1203] The user selects the character they want to interact with from the access portal. The device sends the character information to the server, which then loads the chat AI model (e.g., GPT-3) and voice generation model (e.g., Google Cloud Text-to-Speech) corresponding to the selected character and performs initial setup.
[1204] Conversational Interactions
[1205] Users send text or voice messages from their devices. When a voice message is sent, the device converts the speech to text using Google Cloud Speech-to-Text. The device then sends the text message to the server. The server inputs the message into a chat AI model and generates a reply message. The generated reply message is then input into a voice generation model, which generates voice data. The server then sends this voice data to the user's device, which then plays it back.
[1206] Customization features
[1207] Users can customize the tone of the conversation and the character's reactions in the settings screen, and this information is stored in a database on the server and applied to future interactions.
[1208] Specific examples
[1209] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[1210] User: Texts "How was your day?"
[1211] Terminal: Sends a text message to the server.
[1212] Server: Passes the message to the chat AI model (GPT-3) and generates a response message.
[1213] Chat AI model: Generates the response "Today was a really fun day."
[1214] Speech generation model: Converts the response message into speech data.
[1215] Server: Sends the audio data to the user's device.
[1216] Terminal: Plays the audio data and the user hears the character's response.
[1217] Prompt Sentence Examples
[1218] "Ask Character A, 'How was your day?'"
[1219] The invention allows users to enjoy natural and personalized interactions with fictional characters.
[1220] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1221] Step 1:
[1222] User: Access the device's login screen and enter your username and password.
[1223] Input: Username and Password.
[1224] Output: The authentication information is sent to the device.
[1225] Step 2:
[1226] On the device: The authentication information entered by the user is encrypted and sent to the server.
[1227] Input: Encrypted credentials.
[1228] Output: Authentication information sent to the server.
[1229] Step 3:
[1230] Server: Compares the user information stored in the database, and if authentication is successful, displays the access portal; if authentication fails, generates and returns an error message.
[1231] Input: Encrypted credentials.
[1232] Output: The authentication result (success or failure).
[1233] Step 4:
[1234] User: Select the desired character from the character list displayed on the dashboard.
[1235] Input: The ID of the character selected by the user.
[1236] Output: Character selection information is sent to the device.
[1237] Step 5:
[1238] Device: Sends the selected character information to the server.
[1239] Input: Character ID.
[1240] Output: Character information is sent to the server.
[1241] Step 6:
[1242] Server: Loads the chat AI model and voice generation model corresponding to the selected character and performs initial settings.
[1243] Input: Character ID.
[1244] Output: The chat AI model and speech generation model are initialized.
[1245] Step 7:
[1246] User: Send a text or voice message to the character saying, "How was your day?"
[1247] Input: The user's text or voice message.
[1248] Output: The message is sent to the terminal.
[1249] Step 8:
[1250] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1251] Input: The user's voice message.
[1252] Output: A text message.
[1253] Step 9:
[1254] Terminal: Sends a text message to the server.
[1255] Input: A text message.
[1256] Output: A message is sent to the server.
[1257] Step 10:
[1258] Server: Inputs the received text message into the chat AI model and generates a response message.
[1259] Input: A text message.
[1260] Output: The response message.
[1261] Step 11:
[1262] Chat AI model: Analyzes messages and generates a response message such as "I had a great day today."
[1263] Input: The user's text message.
[1264] Output: The response message.
[1265] Step 12:
[1266] Server: Input the response message into the speech generation model and generate speech data.
[1267] Input: The response message.
[1268] Output: Audio data.
[1269] Step 13:
[1270] Speech generation model: Receives the response message and generates speech data.
[1271] Input: The response message.
[1272] Output: Audio data.
[1273] Step 14:
[1274] Server: Sends the generated audio data to the user's device.
[1275] Input: Audio data.
[1276] Output: The audio data is sent to the device.
[1277] Step 15:
[1278] Terminal: Plays the audio data and the user hears the character's response.
[1279] Input: Audio data.
[1280] Output: Playback of audio data.
[1281] (Application example 1)
[1282] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1283] In virtual environments, users are expected to have a natural experience through interactions with fictional characters. However, conventional systems have made it difficult for users to interact with characters in real time within a virtual environment and receive product explanations and guidance, making it difficult to provide a truly immersive experience. The present invention aims to solve these problems by providing a system that allows users to interact with fictional characters in a virtual environment in a natural and interactive manner, and to receive product explanations and guidance.
[1284] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1285] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for transmitting the generated voice data to the user, and means for reproducing a fictional character in the virtual environment based on a selection by the user and providing guidance or product explanations, thereby enabling the user to have natural conversations with the fictional character in the virtual environment in real time and receive product explanations or guidance.
[1286] "Means for receiving input messages from a user" means a mechanism by which the system receives messages input by a user when the user interacts with a fictional character within a virtual environment.
[1287] A "chat AI model" is an artificial intelligence model built on a natural language processing algorithm used to generate natural response messages in response to user input messages.
[1288] A "voice generation model" is a model used to convert a text response message generated by a chat AI model into voice data.
[1289] A "virtual environment" is a computer-generated virtual space experienced by a user, for example, using a VR headset or smart device.
[1290] A "fictional character" is a fictional character that appears in fictional works such as movies, anime, and games, and whose role is to provide an experience through interactions with users.
[1291] "User authentication" is the process of identifying and verifying a user's identity when accessing a system.
[1292] The "dashboard" is the interface that users can access after authentication, and is the main screen for selecting fictional characters.
[1293] "Real-time response" is a function that generates a response immediately to user input, allowing the character to respond immediately.
[1294] A "database" is a system that stores information such as user settings and conversation history and can read that information as needed.
[1295] "Guide and product information" refers to an explanation or introduction given to a user by a fictional character, providing product details and purchasing advice.
[1296] A specific embodiment of the present invention is to build a system that allows users to naturally interact with fictional characters in a virtual environment. This system is composed of a user terminal and a cloud server.
[1297] System Configuration
[1298] Hardware and Software
[1299] User devices: smartphones (iOS or Android), VR headsets (e.g., Oculus Quest 2), smart glasses.
[1300] Server: Cloud service (e.g. AWS EC2).
[1301] Database: A cloud database service (e.g. AWS RDS).
[1302] AI models: Chat AI models (e.g., OpenAI GPT-4), speech generation models (e.g., Google Wavenet).
[1303] Program processing
[1304] 1. User authentication:
[1305] When a user logs in to the app, the device sends authentication information to the server.
[1306] The server checks the authentication information against the database, and if authentication is successful, displays the dashboard on the terminal.
[1307] 2. Character Selection:
[1308] Users select a fictional character from the dashboard.
[1309] The device sends the selected character information to the server, and the server loads the chat AI model and voice generation model for the character and performs initial settings.
[1310] 3. Conversational Interactions:
[1311] When the user types "Hello, what do you recommend today?", the device sends this message to the server.
[1312] The server passes the input message to the chat AI model to generate a response message.
[1313] The chat AI model generates a response message saying, "Hello! This month's recommendation is new earphones," and returns it to the server.
[1314] The server sends a response message to the speech generation model to generate speech data.
[1315] The generated audio data is returned to the terminal, which then plays the audio.
[1316] 4. Enhanced user experience:
[1317] Users can enjoy a virtual experience in which fictional characters act as guides, showing them around the store and explaining products.
[1318] Conversation history and user preferences are stored in a database for later reuse.
[1319] Specific examples
[1320] For example, if the user selects "Fictional Character B":
[1321] 1. Put on the VR headset.
[1322] 2. A user sends a text message asking, "What are your recommendations in the store?"
[1323] 3. The chat AI model responds, "The recommended product is the newly released smartwatch."
[1324] 4. The speech generation model converts the response into speech.
[1325] 5. Playback on the user's device.
[1326] In this way, it is possible to provide users with an engaging experience by realizing natural interactions and real-time responses with fictional characters within a virtual environment.
[1327] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1328] Step 1:
[1329] User authentication
[1330] The user logs in to the app and enters their user ID and password.
[1331] The terminal transmits the input authentication information to the server.
[1332] The server checks the authentication information against the database, and if authentication is successful, generates dashboard information along with an authentication success message and sends it to the terminal. If authentication fails, it returns an error message.
[1333] The device displays the received dashboard information to the user and waits for the next action.
[1334] Step 2:
[1335] Character Selection
[1336] Users select the fictional character they want from the dashboard.
[1337] The terminal transmits the selected character information to the server and displays a confirmation message.
[1338] The server loads the chat AI model and voice generation model for the character, performs initial setup, and returns a message to the device indicating that the setup is complete.
[1339] The terminal will display an initialization complete message to the user and wait for further action.
[1340] Step 3:
[1341] Interaction Start
[1342] The user types a text message (e.g., "Hello, what do you recommend today?") or a voice message to the selected character.
[1343] If the message is a voice message, the terminal performs voice recognition processing to convert it into text, and then sends the text message to the server.
[1344] The server inputs the received text message into the chat AI model and generates a response message.
[1345] The chat AI model generates a response message such as "Hello! This month's recommendation is new earphones" and returns it to the server.
[1346] Step 4:
[1347] Generating a response message
[1348] The server sends the response message received from the chat AI model to the voice generation model, which generates voice data.
[1349] The speech generation model analyzes the text message and generates corresponding speech data, which is returned to the server.
[1350] Step 5:
[1351] Sending and playing reply messages
[1352] The server transmits the generated voice data to the terminal.
[1353] The device plays back the received audio data, allowing the user to hear the fictional character's response.
[1354] Along with the response message, it displays UI elements to determine the interaction status and next action.
[1355] Step 6:
[1356] Save conversation history and settings
[1357] The server stores conversation history and user settings (e.g. character customization) in a database.
[1358] The database stores the necessary information to allow users to access their conversation history and preferences the next time they log in.
[1359] This series of processes allows users to have natural conversations with fictional characters in a virtual environment, receive responses in real time, and enjoy product guides and explanations in a virtual store.
[1360] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1361] This system recreates fictional characters and allows natural dialogue with users. It also incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly, providing more personalized interactions. The system consists of a server and a user terminal.
[1362] System Overview
[1363] The system's server handles the main processing and responds to user requests using a chat AI model, a voice generation model, and an emotion engine. Users can access the system via their devices and enjoy interacting with the characters.
[1364] User authentication and character selection
[1365] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[1366] 2. Terminal: Sends the entered authentication information to the server.
[1367] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[1368] 4. User: Select the fictional character of your choice from the dashboard.
[1369] 5. Terminal: Sends the selected character information to the server.
[1370] 6. Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[1371] Conversational Interactions
[1372] 1. User: Send a text or voice message to the character from your device.
[1373] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[1374] 3. Device: Sends a text message to the server.
[1375] 4. Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[1376] 5. Emotion Engine: Recognizes user emotions from text messages and generates emotional status.
[1377] 6. Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[1378] 7. Chat AI model: Generates a response message and returns it to the server.
[1379] 8. Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[1380] 9. Speech generation model: Generates speech data based on text data and emotional status.
[1381] 10. Server: Sends the generated voice data to the user's device.
[1382] 11. Terminal: Plays the audio data and the user hears the character's response.
[1383] Specific examples
[1384] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[1385] 1. User: Texts "I'm so tired today."
[1386] 2. Server: Receives the message and passes it to the emotion engine.
[1387] 3. Emotion engine: Analyzes the message and generates an emotional status that identifies the user as tired.
[1388] 4. Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been tough, take a break."
[1389] 5. Speech generation model: Receives a response message and generates speech data in a gentle tone according to the emotional status.
[1390] 6. Server: Sends the audio data to the user's device.
[1391] 7. Terminal: Play the audio data and the character responds in a gentle tone, saying, "That must have been hard. Take a break."
[1392] Customization features
[1393] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[1394] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1395] This system allows users to enjoy natural and emotional interactions with their chosen fictional characters, and individual settings can provide a more personalized experience, allowing for deeper emotional fulfillment for people known as fictiosexuals.
[1396] The processing flow will be explained below.
[1397] This invention is a system that recreates fictional characters and enables natural dialogue with users, and by incorporating an emotion engine that recognizes the user's emotions and adjusts the response content, it provides more personalized interactions. The specific processing flow of the system is explained below in steps.
[1398] User authentication and character selection
[1399] Step 1:
[1400] User: Accesses the system login screen from a terminal and enters a username and password.
[1401] Step 2:
[1402] Terminal: Sends the entered authentication information to the server.
[1403] Step 3:
[1404] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[1405] Step 4:
[1406] Users: Select the fictional character of their choice from the dashboard.
[1407] Step 5:
[1408] Device: Sends the selected character information to the server.
[1409] Step 6:
[1410] Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[1411] Conversational Interactions
[1412] Step 7:
[1413] User: Send a text or voice message from your device to the character.
[1414] Step 8:
[1415] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1416] Step 9:
[1417] Terminal: Sends a text message to the server.
[1418] Step 10:
[1419] Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[1420] Step 11:
[1421] Emotion Engine: Recognizes user emotions from text messages and generates emotional statuses.
[1422] Step 12:
[1423] Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[1424] Step 13:
[1425] Chat AI model: Analyzes messages and generates appropriate response messages.
[1426] Step 14:
[1427] Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[1428] Step 15:
[1429] Speech generation model: Generates speech data based on text data and emotional status.
[1430] Step 16:
[1431] Server: Sends the generated audio data to the user's device.
[1432] Step 17:
[1433] Terminal: Plays the audio data and the user hears the character's response.
[1434] Specific examples
[1435] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[1436] Step 1:
[1437] User: Texts "I'm so tired today."
[1438] Step 2:
[1439] Server: Receives messages and passes them to the emotion engine.
[1440] Step 3:
[1441] Emotion engine: Analyzes messages and generates an emotional status that identifies the user as tired.
[1442] Step 4:
[1443] Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been hard, take a break."
[1444] Step 5:
[1445] Speech generation model: Receives a response message and generates voice data in a gentle tone according to the emotional status.
[1446] Step 6:
[1447] Server: Sends the audio data to the user's device.
[1448] Step 7:
[1449] Device: Plays the audio data and the character responds in a gentle tone, "That must have been hard, take a break."
[1450] Customization features
[1451] Step 1:
[1452] User: You can set preferences to customize the tone of the dialogue and character reactions.
[1453] Step 2:
[1454] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1455] This allows the system to realize natural and emotional interactions with fictional characters selected by the user, making it possible to more deeply satisfy the emotions of people known as fictiosexuals.
[1456] Example 2
[1457] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1458] Conventional dialogue systems have difficulty in engaging in natural conversations with fictional characters, limiting their ability to provide personalized responses based on the user's emotions. Furthermore, the accuracy of speech recognition and speech synthesis is low, making it difficult to generate speech that takes emotions into account. Furthermore, systems lack a mechanism for centrally managing user authentication information and settings and saving dialogue history to provide a better user experience. Therefore, a new dialogue system is needed to improve user satisfaction.
[1459] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1460] In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue generation model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice synthesis model to convert the response message into voice data and generating voice data, means for utilizing an emotion analysis engine to analyze the user's emotion from the input message, means for generating a response message taking into account the emotional status obtained by the emotion analysis engine, means for utilizing the voice synthesis model to generate voice data with a tone corresponding to the emotional status, and means for transmitting the generated voice data to the user. This makes it possible to generate a response message and voice data that take into account the user's emotion and provide a natural and personalized dialogue with a fictional character.
[1461] "User" means any individual or entity that accesses the System and interacts with fictional characters.
[1462] "Input Message" refers to text or voice information sent by a user to a system.
[1463] "Natural language processing algorithms" refer to technologies for understanding a user's input message and generating an appropriate response.
[1464] "Dialogue generation model" refers to a machine learning model that uses a natural language processing algorithm to generate a response message to a user's input message.
[1465] "Response message" refers to text information that the dialogue generation model generates in response to a user's input message.
[1466] "Speech synthesis model" refers to the technology for converting text response messages into speech data.
[1467] "Audio Data" refers to an audio file or audio stream generated by a speech synthesis model.
[1468] "Sentiment analysis engine" refers to technology for analyzing emotions from a user's input message and generating an emotional status.
[1469] "Emotional status" refers to information that expresses a user's emotions numerically or categorically, as determined by the emotion analysis engine.
[1470] "Authentication Information" refers to information such as username and password that a user uses to log into a system.
[1471] "Dashboard" refers to the interface that allows users to select characters and change settings.
[1472] "Fictional Character" refers to any imaginary or fictional person or entity that is recreated within the system.
[1473] "Voice recognition means" refers to technology used to convert a voice message sent by a user into a text message.
[1474] "Conversation history" refers to a record of the dialogue between a user and a character.
[1475] "Database" refers to a management system for systematically storing information necessary for system operation.
[1476] This invention is a system that allows users to have natural and emotional conversations with fictional characters. This system analyzes the user's emotions and appropriately adjusts the content and tone of the conversation to provide a more personalized conversation. Specifically, the system uses the following hardware and software to process data and perform calculations.
[1477] Hardware and software used
[1478] server
[1479] Database Management System (DBMS): Stores user authentication information and configuration data, typically using MySQL or PostgreSQL.
[1480] Dialogue generation model: Generates dialogue. As an example, GPT-3, a machine learning model, is used.
[1481] Speech synthesis model: Converts the response message into audio data. For example, speech synthesis APIs such as Google Cloud Text-to-Speech and IBM Watson Text to Speech are used.
[1482] Sentiment analysis engine: Recognizes user emotions. For example, Microsoft Azure's Sentiment Analysis is used.
[1483] Terminal
[1484] User Interface (UI): A screen where users can interact with fictional characters. Most commonly, it is a web app that is made up of HTML / CSS and JavaScript.
[1485] Speech recognition method: Converts voice input into text. Google Cloud Speech-to-Text and IBM Watson Speech to Text are used.
[1486] User authentication and character selection
[1487] The system begins when a user logs in using their authentication information. The user accesses the system's login screen from their terminal and enters their username and password. The input information is sent from the terminal to the server, which then checks it against a database to authenticate the user.
[1488] After successful authentication, the user is redirected to a dashboard where they can select their desired fictional character. The selected character information is sent from the device to the server, which then loads and initializes the character's corresponding dialogue generation model, speech synthesis model, and sentiment analysis engine.
[1489] Conversational Interactions
[1490] When a user starts a dialogue with a character, an input message is sent from the terminal to the server. If this input message is a voice message, the terminal converts it into text using a voice recognition means.
[1491] The received text message is processed by the server and passed to the sentiment analysis engine, which analyzes the message and generates the user's emotional status. The emotional status is input into the dialogue generation model, and a response message is generated that takes the user's emotions into account.
[1492] The generated response message and emotional status are passed to a speech synthesis model to generate speech data, including tones, which is then sent from the server to the user's device and played back on the device.
[1493] Customization features
[1494] The system provides users with settings to customize the tone of dialogue and character reactions. Users can adjust various settings by manipulating sliders and checkboxes on the settings page. These settings are saved in a database by the server and applied the next time a dialogue is held.
[1495] Specific examples
[1496] For example, if a user sends a message to "Fictional Character A" saying "I'm really tired today," the system operates as follows: First, the user sends a text message saying "I'm really tired today." The server receives this message and passes it to the emotion analysis engine. The emotion analysis engine recognizes that the user is tired and generates an emotion status.
[1497] Next, the server inputs the emotional status into a dialogue generation model and generates a response message that matches the user's emotion, "That must have been hard. Take a break." The speech synthesis model receives this response message and generates voice data in a gentle tone according to the emotional status. Finally, the server sends the voice data to the user's device, which plays it back, with the character responding in a gentle tone, "That must have been hard. Take a break."
[1498] Prompt Sentence Examples
[1499] If a user sends a message to fictional character A saying, "I'm really tired today," please explain in detail how the emotion engine generates the emotion status and what response the dialogue generation model makes.
[1500] As described above, the present invention is a system that allows users to enjoy natural and emotional interactions with fictional characters of their choice, providing a more personalized experience through individual settings.
[1501] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1502] Divide the program's processing flow into processing steps
[1503] Step 1:
[1504] Step 2:
[1505] Step 3:
[1506] ...
[1507] Specific explanations for each step
[1508] Step 1: User authentication and login
[1509] Input: Username, Password
[1510] Output: Authentication result (success or failure), dashboard screen or error message
[1511] Operation:
[1512] 1. User: Enter your username and password on the device login screen.
[1513] 2. Terminal: Sends the input information to the server as a POST request.
[1514] 3. Server: Query the database to verify the authentication information. If authentication is successful, return the dashboard screen; if authentication is unsuccessful, return an error message.
[1515] Step 2: Choose your character
[1516] Input: Character ID
[1517] Output: Character data, initialization result
[1518] Operation:
[1519] 1. User: Select the fictional character of your choice from the dashboard.
[1520] 2. Device: Send the selected character ID to the server as an AJAX request.
[1521] 3. Server: Loads the dialogue generation model, speech synthesis model, and emotion analysis engine corresponding to the character and performs initial configuration.
[1522] Step 3: Send the message
[1523] Input: Text or voice message
[1524] Output: Text message (in case of voice)
[1525] Operation:
[1526] 1. User: Send a text or voice message from your device to the character.
[1527] 2. Device: When a voice message is sent, a speech recognition tool is used to convert the voice into text, the audio file is sent to a cloud API, and the text is returned.
[1528] Step 4: Sentiment Analysis
[1529] Input: Text message
[1530] Output: Emotional status
[1531] Operation:
[1532] 1. Terminal: Sends the converted text message to the server.
[1533] 2. Server: Inputs the received text message into the sentiment analysis engine to analyze the user's sentiment. The sentiment engine analyzes the message and generates a sentiment score.
[1534] Step 5: Generate a response message
[1535] Input: Text message, emotional status
[1536] Output: Response message
[1537] Operation:
[1538] 1. Server: Inputs the emotional status into the dialogue generation model and generates a response message that takes the user's emotions into account. The generative AI model analyzes the text message and emotional status and generates an appropriate response.
[1539] Step 6: Generate audio data
[1540] Input: Response message, emotional status
[1541] Output: Audio data
[1542] Operation:
[1543] 1. Server: Input the response message and emotional status into the speech synthesis model and generate speech data with a tone corresponding to the emotion. The speech synthesis API generates an audio file with a tone based on the emotion.
[1544] Step 7: Playing back audio data
[1545] Input: Audio data
[1546] Output: Play audio
[1547] Operation:
[1548] 1. Server: Sends the generated voice data to the user's device.
[1549] 2. Device: The audio data is played and the user hears the character's response. The device's audio player plays the audio data.
[1550] This is the flow of the program's processing. The specific operations, inputs, and outputs at each step are explained in detail, and it contains all the elements necessary for users to enjoy natural and emotional interactions with fictional characters.
[1551] (Application example 2)
[1552] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1553] Conventional chat systems and voice dialogue systems have difficulty responding to user emotions during interactions, resulting in a limited user experience. In particular, when dealing with customers in physical stores, flexible responses based on customer emotions are required, but the means to achieve this are not fully developed. As a result, there is a risk of customer satisfaction decreasing.
[1554] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue AI model using a natural language processing algorithm for responding to the input message and generating a response message, means for initializing a voice generation algorithm for converting the response message into voice data and generating voice data, means for adding emotion to the response message using an emotion analysis engine that analyzes the user's emotion, and means for transmitting the generated voice data to the user. This enables natural and personalized dialogue that takes the user's emotion into consideration.
[1555] "Means for receiving input messages from a user" means a system or function for obtaining text or voice messages sent by a user from a terminal.
[1556] A "dialogue AI model using natural language processing algorithms" is an algorithm or model that analyzes a user's input message and generates a response accordingly.
[1557] "Speech generation algorithm" refers to an algorithm or technique for converting a text response message into voice data.
[1558] An "emotion analysis engine" is an analysis system that identifies emotions from the user's input message and reflects those emotions in the response message.
[1559] "Means for verifying user authentication information" refers to a system or function that checks authentication information such as a user's ID and password to determine whether the user is legitimate.
[1560] The "means for displaying a dashboard" refers to a system or function for providing an operation screen that is displayed after a user logs in to the system.
[1561] A "means for loading fictional character data" is a system or function that loads data or settings related to a specific fictional character into memory and makes them available for use.
[1562] "Means for analyzing the user's emotional status and reflecting it in the initial settings" refers to a system or function for analyzing the user's emotions and reflecting the results in the character's settings and behavior.
[1563] "Speech recognition means" refers to technology or systems that convert the user's speech into text.
[1564] "Means for storing conversation history and user settings in a database" refers to a system or function that records the content of conversations with users and user settings (such as customization information) in a database, allowing them to be searched or referenced when necessary.
[1565] A "means for customizing fictional character responses" is a system or feature that allows a fictional character's words and actions to be flexibly changed based on the user's settings and choices.
[1566] In this invention, a natural, emotion-conscious dialogue between a user and a fictional character can be realized using the following system configuration and processing procedure.
[1567] Overall system configuration
[1568] The server contains the following main modules:
[1569] 1. Input message receiving module
[1570] Receive text and voice messages from user devices.
[1571] 2. Dialogue AI Model
[1572] It uses natural language processing algorithms to respond to input messages from users.
[1573] 3. Speech Generation Algorithm
[1574] The response message is converted into voice data.
[1575] 4. Sentiment Analysis Engine
[1576] Analyze the user's emotions from the input message and reflect them in the response message.
[1577] 5. User Authentication Module
[1578] Verify the user's credentials and display the dashboard if authentication is successful.
[1579] 6. Character Data Load Module
[1580] Loads a fictional character of the user's choice and performs initial setup.
[1581] 7. Speech Recognition Module
[1582] Convert voice messages to text.
[1583] 8. Database Management Module
[1584] Conversation history and user preferences are stored in a database and loaded as needed.
[1585] 9. Character Reaction Customization Module
[1586] Customize character reactions based on user choices and settings.
[1587] Hardware and Software
[1588] The hardware and software used are as follows:
[1589] Server hardware: A server with a powerful processor and sufficient memory and storage. For example, you can use cloud services such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[1590] Software: Generative AI models such as GPT-3 are used for natural language processing algorithms. WaveNet and Tacotron are used for speech generation algorithms. Sentiment analysis libraries (Python text2emotion, transformers, etc.) are used for sentiment analysis.
[1591] Specific explanation of the process
[1592] server
[1593] 1. The input message receiving module receives a message from the user, which can be text or voice.
[1594] 2. The speech recognition module converts the voice message into text.
[1595] 3. The converted text is input into a sentiment analysis engine to analyze the user's sentiment.
[1596] 4. The analyzed emotional status is input into a dialogue AI model to generate a response message that takes emotions into account.
[1597] 5. The response message is converted into voice data by a voice generation algorithm.
[1598] 6. The generated voice data is sent to the user terminal.
[1599] 7. The database management module stores conversation history and user settings for future use.
[1600] 8. Character reaction customization module modifies character reactions based on user settings.
[1601] Examples and prompts
[1602] For example, if a customer asks "What products do you recommend?" in a physical store, the specific processing is as follows:
[1603] User: Ask "What products do you recommend?"
[1604] Server: The sentiment analysis engine analyzes the sentiment of the question "What product do you recommend?", and the dialogue AI model generates a response message.
[1605] The voice generation algorithm responds, "Today's recommendation is fresh apples," and the generated voice data is sent to the user's device.
[1606] Example prompt sentence:
[1607] "Please enter a user message: What products do you recommend?"
[1608] "Please perform sentiment analysis on this message and generate an appropriate response."
[1609] In this way, a natural and emotional interaction with the user can be achieved.
[1610] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1611] Step 1:
[1612] The user inputs a voice message or text message from a device such as a smartphone or tablet, and the input message is sent from the device to the server.
[1613] Step 2:
[1614] The server invokes a speech recognition module to analyze the received message. When a voice message is sent, the speech recognition module converts the voice data into text data. The text data is used as input and the speech recognition result is obtained as output.
[1615] Step 3:
[1616] The server inputs the text message into the sentiment analysis engine. The sentiment analysis engine analyzes the emotional status from the user's input message. In this step, the sentiment analysis engine analyzes the text data and extracts the emotional status from the user's text message. The input is the text data and the output is the emotional status.
[1617] Step 4:
[1618] The server inputs the analyzed emotional status into a dialogue AI model to generate a response message that takes the user's emotions into account. The dialogue AI model is based on a generative AI model, which receives the emotional status and text message as input and generates a response message as output. This process results in an appropriate response that corresponds to the user's input and emotions.
[1619] Step 5:
[1620] The server sends the generated response message to a voice generation algorithm to generate voice data. The voice generation algorithm receives the response message and the emotional status as input and outputs the voice data in a tone that reflects the emotion.
[1621] Step 6:
[1622] The server transmits the generated voice data to the user's device, which then plays the received voice data and provides the character's response to the user, allowing the user to experience voice responses that reflect their emotions.
[1623] Step 7:
[1624] The server stores the conversation history and user settings in a database. The database management module uses the conversation data and settings information as input and obtains the saved data as output. This data can be used as a reference for future conversations.
[1625] Step 8:
[1626] The user customizes the character's reactions as needed. The server's character reaction customization module receives the user's settings and generates the character's reactions based on them. In this step, the user's input settings are processed and the customized reactions are output.
[1627] The above are the processing steps of the system that realizes this invention, which allows the user to have a natural conversation experience that takes emotions into consideration.
[1628] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1629] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1630] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1631] [Fourth embodiment]
[1632] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1633] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1634] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1635] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1636] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1637] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1638] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1639] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1640] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1641] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1642] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1643] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1644] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1645] The present invention is a system for recreating fictional characters, and is designed to allow users to have natural conversations with the characters. Specific embodiments for implementing this system are as follows.
[1646] System Overview
[1647] This system consists of a server and a user terminal. Users can access the system through their terminal and enjoy interacting with characters. The server is responsible for the main processing and responds to user requests using a chat AI model and a voice generation model.
[1648] User authentication and character selection
[1649] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[1650] 2. Terminal: Sends the entered authentication information to the server.
[1651] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[1652] 4. User: Select the fictional character of your choice from the dashboard.
[1653] 5. Terminal: Sends the selected character information to the server.
[1654] 6. Server: Loads the chat AI model and voice generation model corresponding to the character and performs initial settings.
[1655] Conversational Interactions
[1656] 1. User: Send a text or voice message to the character from your device.
[1657] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[1658] 3. Device: Sends a text message to the server.
[1659] 4. Server: Inputs the received text message into the chat AI model and generates a response message.
[1660] 5. Chat AI model: Generates a response message and returns it to the server.
[1661] 6. Server: Input the response message into the speech generation model and generate speech data.
[1662] 7. Speech generation model: Generates speech data and returns it to the server.
[1663] 8. Server: Sends the generated voice data to the user's device.
[1664] 9. Terminal: Plays the audio data and the user hears the character's response.
[1665] Specific examples
[1666] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[1667] 1. User: Sends a text message saying, "How was your day?"
[1668] 2. Server: Receives messages and passes them to the chat AI model.
[1669] 3. Chat AI model: Analyzes the message and generates a response message such as, "Today was a really fun day."
[1670] 4. Speech generation model: Receives the response message and generates speech data.
[1671] 5. Server: Sends the audio data to the user's device.
[1672] 6. Terminal: Plays the audio data and the user hears the character's response.
[1673] Customization features
[1674] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[1675] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1676] The system allows users to enjoy natural dialogue with the characters they select, and can provide a more personalized experience through individual settings, which can satisfy the feelings of people known as fictiosexuals.
[1677] The processing flow will be explained below.
[1678] Step 1:
[1679] User: Accesses the system login screen and enters a username and password.
[1680] Step 2:
[1681] Terminal: Sends the entered authentication information to the server.
[1682] Step 3:
[1683] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[1684] Step 4:
[1685] Users: Select the fictional character of their choice from the dashboard.
[1686] Step 5:
[1687] Device: Sends the selected character information to the server.
[1688] Step 6:
[1689] Server: Loads the chat AI model and voice generation model corresponding to the characters and performs initial settings.
[1690] Step 7:
[1691] User: Send a text or voice message from your device to the character.
[1692] Step 8:
[1693] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1694] Step 9:
[1695] Terminal: Sends a text message to the server.
[1696] Step 10:
[1697] Server: Inputs the received text message into the chat AI model and generates a response message.
[1698] Step 11:
[1699] Chat AI model: Analyzes messages and generates appropriate response messages.
[1700] Step 12:
[1701] Server: Input the response message into the speech generation model and generate speech data.
[1702] Step 13:
[1703] Speech generation model: Analyzes text data and generates speech data.
[1704] Step 14:
[1705] Server: Sends the generated audio data to the user's device.
[1706] Step 15:
[1707] Terminal: Plays the audio data and the user hears the character's response.
[1708] Step 16:
[1709] User: Set preferences to customize the tone of the dialogue and character reactions.
[1710] Step 17:
[1711] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1712] This series of steps allows users to enjoy natural interactions with their chosen fictional characters.
[1713] Example 1
[1714] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1715] With conventional dialogue systems, it has been difficult to achieve natural and personalized dialogue with fictional characters. It has also been difficult to properly reflect the user's ratings and preferences and respond individually. This has resulted in inconsistent dialogue experiences with the selected characters, leading to low satisfaction. Furthermore, the process of voice message recognition and response generation is complex, making system setup and use time-consuming.
[1716] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1717] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for sending the generated voice data to the user, means for comparing authentication information entered by the user with a database and displaying an access portal if authentication is successful, means for loading data of a virtual character selected by the user and performing initial settings, means for speech recognition to convert voice messages into text, and means for saving conversation history and user settings in a database and loading them as needed, thereby enabling natural and personalized interactions with fictional characters.
[1718] The "means for receiving an input message from a user" refers to a device or program that has the function of transmitting a text message or voice message input by a user to a server.
[1719] "Means for initializing a chat AI model using a natural language processing algorithm and generating a response message" refers to a device or program that has the function of initiating the operation of a chat AI model using natural language processing technology and generating an appropriate response to a message from a user.
[1720] The "means for initializing a speech generation model and generating speech data" refers to a device or program having the function of operating a speech generation model for converting a text message into speech data and converting the generated text response into speech data.
[1721] The "means for transmitting generated voice data to a user" refers to a device or program having the function of transmitting voice data generated by the server to a user's terminal.
[1722] "Means for checking the authentication information entered by the user against a database and displaying an access portal if authentication is successful" refers to a device or program that has the function of checking the authentication information entered by the user (such as a user name and password) against a database to perform authentication, and displaying the main screen if authentication is successful.
[1723] "Means for loading the data of a virtual character selected by the user and performing initial settings" refers to a device or program that has the function of loading the data of a character selected by the user and setting the corresponding chat AI model and voice generation model.
[1724] The "voice recognition means for converting a voice message into text" is a device or program that has the function of converting a voice message uttered by a user into text format.
[1725] "Means for saving conversation history and user settings in a database and loading them as needed" refers to a device or program that has the function of recording the conversation history and user settings in a database and referencing this information during subsequent conversations.
[1726] This invention relates to a system that allows users to interact with fictional characters. The system consists of a server and a user terminal, and utilizes a chat AI model and a voice generation model to realize natural and personalized interactions.
[1727] System configuration
[1728] The system mainly consists of the following components:
[1729] 1. Server: A high-performance computer is used to process the chat AI model and voice generation model.
[1730] 2. Terminal: The device through which the user accesses the service, including smartphones and PCs.
[1731] 3. Database: Stores authentication information, character settings, conversation history, and user preferences.
[1732] Hardware and software used
[1733] Hardware: High-performance servers (e.g., cloud service provider servers), user devices (smartphones, PCs)
[1734] Software: Database (e.g., MySQL), chat AI model (e.g., GPT-3), speech generation model (e.g., Google Cloud Text-to-Speech), speech recognition method (e.g., Google Cloud Speech-to-Text)
[1735] Explanation of program processing
[1736] User authentication
[1737] The user accesses the system's login screen from a terminal and enters their username and password. The terminal encrypts this authentication information and sends it to the server, which checks it against a database and, if authentication is successful, displays the access portal. This process uses a MySQL database.
[1738] Character Selection
[1739] The user selects the character they want to interact with from the access portal. The device sends the character information to the server, which then loads the chat AI model (e.g., GPT-3) and voice generation model (e.g., Google Cloud Text-to-Speech) corresponding to the selected character and performs initial setup.
[1740] Conversational Interactions
[1741] Users send text or voice messages from their devices. When a voice message is sent, the device converts the speech to text using Google Cloud Speech-to-Text. The device then sends the text message to the server. The server inputs the message into a chat AI model and generates a reply message. The generated reply message is then input into a voice generation model, which generates voice data. The server then sends this voice data to the user's device, which then plays it back.
[1742] Customization features
[1743] Users can customize the tone of the conversation and the character's reactions in the settings screen, and this information is stored in a database on the server and applied to future interactions.
[1744] Specific examples
[1745] For example, if the user selects "Fictional Character A," the dialogue will proceed as follows:
[1746] User: Texts "How was your day?"
[1747] Terminal: Sends a text message to the server.
[1748] Server: Passes the message to the chat AI model (GPT-3) and generates a response message.
[1749] Chat AI model: Generates the response "Today was a really fun day."
[1750] Speech generation model: Converts the response message into speech data.
[1751] Server: Sends the audio data to the user's device.
[1752] Terminal: Plays the audio data and the user hears the character's response.
[1753] Prompt Sentence Examples
[1754] "Ask Character A, 'How was your day?'"
[1755] The invention allows users to enjoy natural and personalized interactions with fictional characters.
[1756] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1757] Step 1:
[1758] User: Access the device's login screen and enter your username and password.
[1759] Input: Username and Password.
[1760] Output: The authentication information is sent to the device.
[1761] Step 2:
[1762] On the device: The authentication information entered by the user is encrypted and sent to the server.
[1763] Input: Encrypted credentials.
[1764] Output: Authentication information sent to the server.
[1765] Step 3:
[1766] Server: Compares the user information stored in the database, and if authentication is successful, displays the access portal; if authentication fails, generates and returns an error message.
[1767] Input: Encrypted credentials.
[1768] Output: The authentication result (success or failure).
[1769] Step 4:
[1770] User: Select the desired character from the character list displayed on the dashboard.
[1771] Input: The ID of the character selected by the user.
[1772] Output: Character selection information is sent to the device.
[1773] Step 5:
[1774] Device: Sends the selected character information to the server.
[1775] Input: Character ID.
[1776] Output: Character information is sent to the server.
[1777] Step 6:
[1778] Server: Loads the chat AI model and voice generation model corresponding to the selected character and performs initial settings.
[1779] Input: Character ID.
[1780] Output: The chat AI model and speech generation model are initialized.
[1781] Step 7:
[1782] User: Send a text or voice message to the character saying, "How was your day?"
[1783] Input: The user's text or voice message.
[1784] Output: The message is sent to the terminal.
[1785] Step 8:
[1786] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1787] Input: The user's voice message.
[1788] Output: A text message.
[1789] Step 9:
[1790] Terminal: Sends a text message to the server.
[1791] Input: A text message.
[1792] Output: A message is sent to the server.
[1793] Step 10:
[1794] Server: Inputs the received text message into the chat AI model and generates a response message.
[1795] Input: A text message.
[1796] Output: The response message.
[1797] Step 11:
[1798] Chat AI model: Analyzes messages and generates a response message such as "I had a great day today."
[1799] Input: The user's text message.
[1800] Output: The response message.
[1801] Step 12:
[1802] Server: Input the response message into the speech generation model and generate speech data.
[1803] Input: The response message.
[1804] Output: Audio data.
[1805] Step 13:
[1806] Speech generation model: Receives the response message and generates speech data.
[1807] Input: The response message.
[1808] Output: Audio data.
[1809] Step 14:
[1810] Server: Sends the generated audio data to the user's device.
[1811] Input: Audio data.
[1812] Output: The audio data is sent to the device.
[1813] Step 15:
[1814] Terminal: Plays the audio data and the user hears the character's response.
[1815] Input: Audio data.
[1816] Output: Playback of audio data.
[1817] (Application example 1)
[1818] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1819] In virtual environments, users are expected to have a natural experience through interactions with fictional characters. However, conventional systems have made it difficult for users to interact with characters in real time within a virtual environment and receive product explanations and guidance, making it difficult to provide a truly immersive experience. The present invention aims to solve these problems by providing a system that allows users to interact with fictional characters in a virtual environment in a natural and interactive manner, and to receive product explanations and guidance.
[1820] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1821] In this invention, the server includes means for receiving an input message from a user, means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice generation model to convert the response message into voice data and generating the voice data, means for transmitting the generated voice data to the user, and means for reproducing a fictional character in the virtual environment based on a selection by the user and providing guidance or product explanations, thereby enabling the user to have natural conversations with the fictional character in the virtual environment in real time and receive product explanations or guidance.
[1822] "Means for receiving input messages from a user" means a mechanism by which the system receives messages input by a user when the user interacts with a fictional character within a virtual environment.
[1823] A "chat AI model" is an artificial intelligence model built on a natural language processing algorithm used to generate natural response messages in response to user input messages.
[1824] A "voice generation model" is a model used to convert a text response message generated by a chat AI model into voice data.
[1825] A "virtual environment" is a computer-generated virtual space experienced by a user, for example, using a VR headset or smart device.
[1826] A "fictional character" is a fictional character that appears in fictional works such as movies, anime, and games, and whose role is to provide an experience through interactions with users.
[1827] "User authentication" is the process of identifying and verifying a user's identity when accessing a system.
[1828] The "dashboard" is the interface that users can access after authentication, and is the main screen for selecting fictional characters.
[1829] "Real-time response" is a function that generates a response immediately to user input, allowing the character to respond immediately.
[1830] A "database" is a system that stores information such as user settings and conversation history and can read that information as needed.
[1831] "Guide and product information" refers to an explanation or introduction given to a user by a fictional character, providing product details and purchasing advice.
[1832] A specific embodiment of the present invention is to build a system that allows users to naturally interact with fictional characters in a virtual environment. This system is composed of a user terminal and a cloud server.
[1833] System Configuration
[1834] Hardware and Software
[1835] User devices: smartphones (iOS or Android), VR headsets (e.g., Oculus Quest 2), smart glasses.
[1836] Server: Cloud service (e.g. AWS EC2).
[1837] Database: A cloud database service (e.g. AWS RDS).
[1838] AI models: Chat AI models (e.g., OpenAI GPT-4), speech generation models (e.g., Google Wavenet).
[1839] Program processing
[1840] 1. User authentication:
[1841] When a user logs in to the app, the device sends authentication information to the server.
[1842] The server checks the authentication information against the database, and if authentication is successful, displays the dashboard on the terminal.
[1843] 2. Character Selection:
[1844] Users select a fictional character from the dashboard.
[1845] The device sends the selected character information to the server, and the server loads the chat AI model and voice generation model for the character and performs initial settings.
[1846] 3. Conversational Interactions:
[1847] When the user types "Hello, what do you recommend today?", the device sends this message to the server.
[1848] The server passes the input message to the chat AI model to generate a response message.
[1849] The chat AI model generates a response message saying, "Hello! This month's recommendation is new earphones," and returns it to the server.
[1850] The server sends a response message to the speech generation model to generate speech data.
[1851] The generated audio data is returned to the terminal, which then plays the audio.
[1852] 4. Enhanced user experience:
[1853] Users can enjoy a virtual experience in which fictional characters act as guides, showing them around the store and explaining products.
[1854] Conversation history and user preferences are stored in a database for later reuse.
[1855] Specific examples
[1856] For example, if the user selects "Fictional Character B":
[1857] 1. Put on the VR headset.
[1858] 2. A user sends a text message asking, "What are your recommendations in the store?"
[1859] 3. The chat AI model responds, "The recommended product is the newly released smartwatch."
[1860] 4. The speech generation model converts the response into speech.
[1861] 5. Playback on the user's device.
[1862] In this way, it is possible to provide users with an engaging experience by realizing natural interactions and real-time responses with fictional characters within a virtual environment.
[1863] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1864] Step 1:
[1865] User authentication
[1866] The user logs in to the app and enters their user ID and password.
[1867] The terminal transmits the input authentication information to the server.
[1868] The server checks the authentication information against the database, and if authentication is successful, generates dashboard information along with an authentication success message and sends it to the terminal. If authentication fails, it returns an error message.
[1869] The device displays the received dashboard information to the user and waits for the next action.
[1870] Step 2:
[1871] Character Selection
[1872] Users select the fictional character they want from the dashboard.
[1873] The terminal transmits the selected character information to the server and displays a confirmation message.
[1874] The server loads the chat AI model and voice generation model for the character, performs initial setup, and returns a message to the device indicating that the setup is complete.
[1875] The terminal will display an initialization complete message to the user and wait for further action.
[1876] Step 3:
[1877] Interaction Start
[1878] The user types a text message (e.g., "Hello, what do you recommend today?") or a voice message to the selected character.
[1879] If the message is a voice message, the terminal performs voice recognition processing to convert it into text, and then sends the text message to the server.
[1880] The server inputs the received text message into the chat AI model and generates a response message.
[1881] The chat AI model generates a response message such as "Hello! This month's recommendation is new earphones" and returns it to the server.
[1882] Step 4:
[1883] Generating a response message
[1884] The server sends the response message received from the chat AI model to the voice generation model, which generates voice data.
[1885] The speech generation model analyzes the text message and generates corresponding speech data, which is returned to the server.
[1886] Step 5:
[1887] Sending and playing reply messages
[1888] The server transmits the generated voice data to the terminal.
[1889] The device plays back the received audio data, allowing the user to hear the fictional character's response.
[1890] Along with the response message, it displays UI elements to determine the interaction status and next action.
[1891] Step 6:
[1892] Save conversation history and settings
[1893] The server stores conversation history and user settings (e.g. character customization) in a database.
[1894] The database stores the necessary information to allow users to access their conversation history and preferences the next time they log in.
[1895] This series of processes allows users to have natural conversations with fictional characters in a virtual environment, receive responses in real time, and enjoy product guides and explanations in a virtual store.
[1896] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1897] This system recreates fictional characters and allows natural dialogue with users. It also incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly, providing more personalized interactions. The system consists of a server and a user terminal.
[1898] System Overview
[1899] The system's server handles the main processing and responds to user requests using a chat AI model, a voice generation model, and an emotion engine. Users can access the system via their devices and enjoy interacting with the characters.
[1900] User authentication and character selection
[1901] 1. User: Accesses the system login screen from a terminal and enters a username and password.
[1902] 2. Terminal: Sends the entered authentication information to the server.
[1903] 3. Server: Checks the user's credentials against the database and displays the dashboard if authentication is successful, otherwise returns an error message.
[1904] 4. User: Select the fictional character of your choice from the dashboard.
[1905] 5. Terminal: Sends the selected character information to the server.
[1906] 6. Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[1907] Conversational Interactions
[1908] 1. User: Send a text or voice message to the character from your device.
[1909] 2. Terminal: When a voice message is sent, the voice is converted into text using a speech recognition means.
[1910] 3. Device: Sends a text message to the server.
[1911] 4. Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[1912] 5. Emotion Engine: Recognizes user emotions from text messages and generates emotional status.
[1913] 6. Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[1914] 7. Chat AI model: Generates a response message and returns it to the server.
[1915] 8. Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[1916] 9. Speech generation model: Generates speech data based on text data and emotional status.
[1917] 10. Server: Sends the generated voice data to the user's device.
[1918] 11. Terminal: Plays the audio data and the user hears the character's response.
[1919] Specific examples
[1920] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[1921] 1. User: Texts "I'm so tired today."
[1922] 2. Server: Receives the message and passes it to the emotion engine.
[1923] 3. Emotion engine: Analyzes the message and generates an emotional status that identifies the user as tired.
[1924] 4. Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been tough, take a break."
[1925] 5. Speech generation model: Receives a response message and generates speech data in a gentle tone according to the emotional status.
[1926] 6. Server: Sends the audio data to the user's device.
[1927] 7. Terminal: Play the audio data and the character responds in a gentle tone, saying, "That must have been hard. Take a break."
[1928] Customization features
[1929] 1. User: You can set preferences to customize the tone of the conversation and character reactions.
[1930] 2. Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1931] This system allows users to enjoy natural and emotional interactions with their chosen fictional characters, and individual settings can provide a more personalized experience, allowing for deeper emotional fulfillment for people known as fictiosexuals.
[1932] The processing flow will be explained below.
[1933] This invention is a system that recreates fictional characters and enables natural dialogue with users, and by incorporating an emotion engine that recognizes the user's emotions and adjusts the response content, it provides more personalized interactions. The specific processing flow of the system is explained below in steps.
[1934] User authentication and character selection
[1935] Step 1:
[1936] User: Accesses the system login screen from a terminal and enters a username and password.
[1937] Step 2:
[1938] Terminal: Sends the entered authentication information to the server.
[1939] Step 3:
[1940] Server: Checks the user's credentials against a database and displays the dashboard if authentication is successful, or returns an error message if authentication fails.
[1941] Step 4:
[1942] Users: Select the fictional character of their choice from the dashboard.
[1943] Step 5:
[1944] Device: Sends the selected character information to the server.
[1945] Step 6:
[1946] Server: Loads the chat AI model, voice generation model, and emotion engine corresponding to the characters and performs initial settings.
[1947] Conversational Interactions
[1948] Step 7:
[1949] User: Send a text or voice message from your device to the character.
[1950] Step 8:
[1951] Terminal: When a voice message is sent, the speech is converted to text using speech recognition means.
[1952] Step 9:
[1953] Terminal: Sends a text message to the server.
[1954] Step 10:
[1955] Server: Inputs the received text message into the emotion engine and analyzes the user's emotions.
[1956] Step 11:
[1957] Emotion Engine: Recognizes user emotions from text messages and generates emotional statuses.
[1958] Step 12:
[1959] Server: Inputs the emotional status into the chat AI model and generates a response message that takes the user's emotions into account.
[1960] Step 13:
[1961] Chat AI model: Analyzes messages and generates appropriate response messages.
[1962] Step 14:
[1963] Server: Input the response message and emotional status into the speech generation model, and generate speech data with a tone appropriate to the emotion.
[1964] Step 15:
[1965] Speech generation model: Generates speech data based on text data and emotional status.
[1966] Step 16:
[1967] Server: Sends the generated audio data to the user's device.
[1968] Step 17:
[1969] Terminal: Plays the audio data and the user hears the character's response.
[1970] Specific examples
[1971] For example, when a user sends a message to "fictional character A" saying "I'm really tired today," the specific processing flow is as follows.
[1972] Step 1:
[1973] User: Texts "I'm so tired today."
[1974] Step 2:
[1975] Server: Receives messages and passes them to the emotion engine.
[1976] Step 3:
[1977] Emotion engine: Analyzes messages and generates an emotional status that identifies the user as tired.
[1978] Step 4:
[1979] Server: Input the emotional status into the chat AI model and generate a response message that matches the user's emotions, such as "That must have been hard, take a break."
[1980] Step 5:
[1981] Speech generation model: Receives a response message and generates voice data in a gentle tone according to the emotional status.
[1982] Step 6:
[1983] Server: Sends the audio data to the user's device.
[1984] Step 7:
[1985] Device: Plays the audio data and the character responds in a gentle tone, "That must have been hard, take a break."
[1986] Customization features
[1987] Step 1:
[1988] User: You can set preferences to customize the tone of the dialogue and character reactions.
[1989] Step 2:
[1990] Server: Saves the user's preferences in a database and applies them the next time the user interacts with the server.
[1991] This allows the system to realize natural and emotional interactions with fictional characters selected by the user, making it possible to more deeply satisfy the emotions of people known as fictiosexuals.
[1992] Example 2
[1993] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1994] Conventional dialogue systems have difficulty in engaging in natural conversations with fictional characters, limiting their ability to provide personalized responses based on the user's emotions. Furthermore, the accuracy of speech recognition and speech synthesis is low, making it difficult to generate speech that takes emotions into account. Furthermore, systems lack a mechanism for centrally managing user authentication information and settings and saving dialogue history to provide a better user experience. Therefore, a new dialogue system is needed to improve user satisfaction.
[1995] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1996] In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue generation model using a natural language processing algorithm to respond to the input message and generating a response message, means for initializing a voice synthesis model to convert the response message into voice data and generating voice data, means for utilizing an emotion analysis engine to analyze the user's emotion from the input message, means for generating a response message taking into account the emotional status obtained by the emotion analysis engine, means for utilizing the voice synthesis model to generate voice data with a tone corresponding to the emotional status, and means for transmitting the generated voice data to the user. This makes it possible to generate a response message and voice data that take into account the user's emotion and provide a natural and personalized dialogue with a fictional character.
[1997] "User" means any individual or entity that accesses the System and interacts with fictional characters.
[1998] "Input Message" refers to text or voice information sent by a user to a system.
[1999] "Natural language processing algorithms" refer to technologies for understanding a user's input message and generating an appropriate response.
[2000] "Dialogue generation model" refers to a machine learning model that uses a natural language processing algorithm to generate a response message to a user's input message.
[2001] "Response message" refers to text information that the dialogue generation model generates in response to a user's input message.
[2002] "Speech synthesis model" refers to the technology for converting text response messages into speech data.
[2003] "Audio Data" refers to an audio file or audio stream generated by a speech synthesis model.
[2004] "Sentiment analysis engine" refers to technology for analyzing emotions from a user's input message and generating an emotional status.
[2005] "Emotional status" refers to information that expresses a user's emotions numerically or categorically, as determined by the emotion analysis engine.
[2006] "Authentication Information" refers to information such as username and password that a user uses to log into a system.
[2007] "Dashboard" refers to the interface that allows users to select characters and change settings.
[2008] "Fictional Character" refers to any imaginary or fictional person or entity that is recreated within the system.
[2009] "Voice recognition means" refers to technology used to convert a voice message sent by a user into a text message.
[2010] "Conversation history" refers to a record of the dialogue between a user and a character.
[2011] "Database" refers to a management system for systematically storing information necessary for system operation.
[2012] This invention is a system that allows users to have natural and emotional conversations with fictional characters. This system analyzes the user's emotions and appropriately adjusts the content and tone of the conversation to provide a more personalized conversation. Specifically, the system uses the following hardware and software to process data and perform calculations.
[2013] Hardware and software used
[2014] server
[2015] Database Management System (DBMS): Stores user authentication information and configuration data, typically using MySQL or PostgreSQL.
[2016] Dialogue generation model: Generates dialogue. As an example, GPT-3, a machine learning model, is used.
[2017] Speech synthesis model: Converts the response message into audio data. For example, speech synthesis APIs such as Google Cloud Text-to-Speech and IBM Watson Text to Speech are used.
[2018] Sentiment analysis engine: Recognizes user emotions. For example, Microsoft Azure's Sentiment Analysis is used.
[2019] Terminal
[2020] User Interface (UI): A screen where users can interact with fictional characters. Most commonly, it is a web app that is made up of HTML / CSS and JavaScript.
[2021] Speech recognition method: Converts voice input into text. Google Cloud Speech-to-Text and IBM Watson Speech to Text are used.
[2022] User authentication and character selection
[2023] The system begins when a user logs in using their authentication information. The user accesses the system's login screen from their terminal and enters their username and password. The input information is sent from the terminal to the server, which then checks it against a database to authenticate the user.
[2024] After successful authentication, the user is redirected to a dashboard where they can select their desired fictional character. The selected character information is sent from the device to the server, which then loads and initializes the character's corresponding dialogue generation model, speech synthesis model, and sentiment analysis engine.
[2025] Conversational Interactions
[2026] When a user starts a dialogue with a character, an input message is sent from the terminal to the server. If this input message is a voice message, the terminal converts it into text using a voice recognition means.
[2027] The received text message is processed by the server and passed to the sentiment analysis engine, which analyzes the message and generates the user's emotional status. The emotional status is input into the dialogue generation model, and a response message is generated that takes the user's emotions into account.
[2028] The generated response message and emotional status are passed to a speech synthesis model to generate speech data, including tones, which is then sent from the server to the user's device and played back on the device.
[2029] Customization features
[2030] The system provides users with settings to customize the tone of dialogue and character reactions. Users can adjust various settings by manipulating sliders and checkboxes on the settings page. These settings are saved in a database by the server and applied the next time a dialogue is held.
[2031] Specific examples
[2032] For example, if a user sends a message to "Fictional Character A" saying "I'm really tired today," the system operates as follows: First, the user sends a text message saying "I'm really tired today." The server receives this message and passes it to the emotion analysis engine. The emotion analysis engine recognizes that the user is tired and generates an emotion status.
[2033] Next, the server inputs the emotional status into a dialogue generation model and generates a response message that matches the user's emotion, "That must have been hard. Take a break." The speech synthesis model receives this response message and generates voice data in a gentle tone according to the emotional status. Finally, the server sends the voice data to the user's device, which plays it back, with the character responding in a gentle tone, "That must have been hard. Take a break."
[2034] Prompt Sentence Examples
[2035] If a user sends a message to fictional character A saying, "I'm really tired today," please explain in detail how the emotion engine generates the emotion status and what response the dialogue generation model makes.
[2036] As described above, the present invention is a system that allows users to enjoy natural and emotional interactions with fictional characters of their choice, providing a more personalized experience through individual settings.
[2037] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2038] Divide the program's processing flow into processing steps
[2039] Step 1:
[2040] Step 2:
[2041] Step 3:
[2042] ...
[2043] Specific explanations for each step
[2044] Step 1: User authentication and login
[2045] Input: Username, Password
[2046] Output: Authentication result (success or failure), dashboard screen or error message
[2047] Operation:
[2048] 1. User: Enter your username and password on the device login screen.
[2049] 2. Terminal: Sends the input information to the server as a POST request.
[2050] 3. Server: Query the database to verify the authentication information. If authentication is successful, return the dashboard screen; if authentication is unsuccessful, return an error message.
[2051] Step 2: Choose your character
[2052] Input: Character ID
[2053] Output: Character data, initialization result
[2054] Operation:
[2055] 1. User: Select the fictional character of your choice from the dashboard.
[2056] 2. Device: Send the selected character ID to the server as an AJAX request.
[2057] 3. Server: Loads the dialogue generation model, speech synthesis model, and emotion analysis engine corresponding to the character and performs initial configuration.
[2058] Step 3: Send the message
[2059] Input: Text or voice message
[2060] Output: Text message (in case of voice)
[2061] Operation:
[2062] 1. User: Send a text or voice message from your device to the character.
[2063] 2. Device: When a voice message is sent, a speech recognition tool is used to convert the voice into text, the audio file is sent to a cloud API, and the text is returned.
[2064] Step 4: Sentiment Analysis
[2065] Input: Text message
[2066] Output: Emotional status
[2067] Operation:
[2068] 1. Terminal: Sends the converted text message to the server.
[2069] 2. Server: Inputs the received text message into the sentiment analysis engine to analyze the user's sentiment. The sentiment engine analyzes the message and generates a sentiment score.
[2070] Step 5: Generate a response message
[2071] Input: Text message, emotional status
[2072] Output: Response message
[2073] Operation:
[2074] 1. Server: Inputs the emotional status into the dialogue generation model and generates a response message that takes the user's emotions into account. The generative AI model analyzes the text message and emotional status and generates an appropriate response.
[2075] Step 6: Generate audio data
[2076] Input: Response message, emotional status
[2077] Output: Audio data
[2078] Operation:
[2079] 1. Server: Input the response message and emotional status into the speech synthesis model and generate speech data with a tone corresponding to the emotion. The speech synthesis API generates an audio file with a tone based on the emotion.
[2080] Step 7: Playing back audio data
[2081] Input: Audio data
[2082] Output: Play audio
[2083] Operation:
[2084] 1. Server: Sends the generated voice data to the user's device.
[2085] 2. Device: The audio data is played and the user hears the character's response. The device's audio player plays the audio data.
[2086] This is the flow of the program's processing. The specific operations, inputs, and outputs at each step are explained in detail, and it contains all the elements necessary for users to enjoy natural and emotional interactions with fictional characters.
[2087] (Application example 2)
[2088] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2089] Conventional chat systems and voice dialogue systems have difficulty responding to user emotions during interactions, resulting in a limited user experience. In particular, when dealing with customers in physical stores, flexible responses based on customer emotions are required, but the means to achieve this are not fully developed. As a result, there is a risk of customer satisfaction decreasing.
[2090] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an input message from a user, means for initializing a dialogue AI model using a natural language processing algorithm for responding to the input message and generating a response message, means for initializing a voice generation algorithm for converting the response message into voice data and generating voice data, means for adding emotion to the response message using an emotion analysis engine that analyzes the user's emotion, and means for transmitting the generated voice data to the user. This enables natural and personalized dialogue that takes the user's emotion into consideration.
[2091] "Means for receiving input messages from a user" means a system or function for obtaining text or voice messages sent by a user from a terminal.
[2092] A "dialogue AI model using natural language processing algorithms" is an algorithm or model that analyzes a user's input message and generates a response accordingly.
[2093] "Speech generation algorithm" refers to an algorithm or technique for converting a text response message into voice data.
[2094] An "emotion analysis engine" is an analysis system that identifies emotions from the user's input message and reflects those emotions in the response message.
[2095] "Means for verifying user authentication information" refers to a system or function that checks authentication information such as a user's ID and password to determine whether the user is legitimate.
[2096] The "means for displaying a dashboard" refers to a system or function for providing an operation screen that is displayed after a user logs in to the system.
[2097] A "means for loading fictional character data" is a system or function that loads data or settings related to a specific fictional character into memory and makes them available for use.
[2098] "Means for analyzing the user's emotional status and reflecting it in the initial settings" refers to a system or function for analyzing the user's emotions and reflecting the results in the character's settings and behavior.
[2099] "Speech recognition means" refers to technology or systems that convert the user's speech into text.
[2100] "Means for storing conversation history and user settings in a database" refers to a system or function that records the content of conversations with users and user settings (such as customization information) in a database, allowing them to be searched or referenced when necessary.
[2101] A "means for customizing fictional character responses" is a system or feature that allows a fictional character's words and actions to be flexibly changed based on the user's settings and choices.
[2102] In this invention, a natural, emotion-conscious dialogue between a user and a fictional character can be realized using the following system configuration and processing procedure.
[2103] Overall system configuration
[2104] The server contains the following main modules:
[2105] 1. Input message receiving module
[2106] Receive text and voice messages from user devices.
[2107] 2. Dialogue AI Model
[2108] It uses natural language processing algorithms to respond to input messages from users.
[2109] 3. Speech Generation Algorithm
[2110] The response message is converted into voice data.
[2111] 4. Sentiment Analysis Engine
[2112] Analyze the user's emotions from the input message and reflect them in the response message.
[2113] 5. User Authentication Module
[2114] Verify the user's credentials and display the dashboard if authentication is successful.
[2115] 6. Character Data Load Module
[2116] Loads a fictional character of the user's choice and performs initial setup.
[2117] 7. Speech Recognition Module
[2118] Convert voice messages to text.
[2119] 8. Database Management Module
[2120] Conversation history and user preferences are stored in a database and loaded as needed.
[2121] 9. Character Reaction Customization Module
[2122] Customize character reactions based on user choices and settings.
[2123] Hardware and Software
[2124] The hardware and software used are as follows:
[2125] Server hardware: A server with a powerful processor and sufficient memory and storage. For example, you can use cloud services such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[2126] Software: Generative AI models such as GPT-3 are used for natural language processing algorithms. WaveNet and Tacotron are used for speech generation algorithms. Sentiment analysis libraries (Python text2emotion, transformers, etc.) are used for sentiment analysis.
[2127] Specific explanation of the process
[2128] server
[2129] 1. The input message receiving module receives a message from the user, which can be text or voice.
[2130] 2. The speech recognition module converts the voice message into text.
[2131] 3. The converted text is input into a sentiment analysis engine to analyze the user's sentiment.
[2132] 4. The analyzed emotional status is input into a dialogue AI model to generate a response message that takes emotions into account.
[2133] 5. The response message is converted into voice data by a voice generation algorithm.
[2134] 6. The generated voice data is sent to the user terminal.
[2135] 7. The database management module stores conversation history and user settings for future use.
[2136] 8. Character reaction customization module modifies character reactions based on user settings.
[2137] Examples and prompts
[2138] For example, if a customer asks "What products do you recommend?" in a physical store, the specific processing is as follows:
[2139] User: Ask "What products do you recommend?"
[2140] Server: The sentiment analysis engine analyzes the sentiment of the question "What product do you recommend?", and the dialogue AI model generates a response message.
[2141] The voice generation algorithm responds, "Today's recommendation is fresh apples," and the generated voice data is sent to the user's device.
[2142] Example prompt sentence:
[2143] "Please enter a user message: What products do you recommend?"
[2144] "Please perform sentiment analysis on this message and generate an appropriate response."
[2145] In this way, a natural and emotional interaction with the user can be achieved.
[2146] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2147] Step 1:
[2148] The user inputs a voice message or text message from a device such as a smartphone or tablet, and the input message is sent from the device to the server.
[2149] Step 2:
[2150] The server invokes a speech recognition module to analyze the received message. When a voice message is sent, the speech recognition module converts the voice data into text data. The text data is used as input and the speech recognition result is obtained as output.
[2151] Step 3:
[2152] The server inputs the text message into the sentiment analysis engine. The sentiment analysis engine analyzes the emotional status from the user's input message. In this step, the sentiment analysis engine analyzes the text data and extracts the emotional status from the user's text message. The input is the text data and the output is the emotional status.
[2153] Step 4:
[2154] The server inputs the analyzed emotional status into a dialogue AI model to generate a response message that takes the user's emotions into account. The dialogue AI model is based on a generative AI model, which receives the emotional status and text message as input and generates a response message as output. This process results in an appropriate response that corresponds to the user's input and emotions.
[2155] Step 5:
[2156] The server sends the generated response message to a voice generation algorithm to generate voice data. The voice generation algorithm receives the response message and the emotional status as input and outputs the voice data in a tone that reflects the emotion.
[2157] Step 6:
[2158] The server transmits the generated voice data to the user's device, which then plays the received voice data and provides the character's response to the user, allowing the user to experience voice responses that reflect their emotions.
[2159] Step 7:
[2160] The server stores the conversation history and user settings in a database. The database management module uses the conversation data and settings information as input and obtains the saved data as output. This data can be used as a reference for future conversations.
[2161] Step 8:
[2162] The user customizes the character's reactions as needed. The server's character reaction customization module receives the user's settings and generates the character's reactions based on them. In this step, the user's input settings are processed and the customized reactions are output.
[2163] The above are the processing steps of the system that realizes this invention, which allows the user to have a natural conversation experience that takes emotions into consideration.
[2164] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2165] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2166] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2167] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2168] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2169] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2170] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2171] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2172] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2173] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2174] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2175] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2176] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2177] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2178] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2179] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2180] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2181] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2182] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2183] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2184] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2185] The following is further disclosed regarding the above embodiment.
[2186] (Claim 1)
[2187] means for receiving input messages from a user;
[2188] A means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message;
[2189] means for initializing a voice generation model for converting the response message into voice data and generating voice data;
[2190] means for transmitting the generated audio data to a user;
[2191] A system including:
[2192] (Claim 2)
[2193] A means to verify the user's credentials and display the dashboard if authentication is successful;
[2194] A means for loading and initializing the data of a fictional character selected by the user;
[2195] The system of claim 1 further comprising:
[2196] (Claim 3)
[2197] a speech recognition means for converting a user's voice message into a text message;
[2198] A means to store conversation history and user preferences in a database and load them as needed;
[2199] The system of claim 1 further comprising:
[2200] "Example 1"
[2201] (Claim 1)
[2202] means for receiving input messages from a user;
[2203] A means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message;
[2204] means for initializing a voice generation model for converting the response message into voice data and generating voice data;
[2205] means for transmitting the generated audio data to a user;
[2206] a means for verifying the authentication information entered by the user against a database and, if authentication is successful, displaying an access portal;
[2207] A means for loading the data of a virtual character selected by the user and performing initial settings;
[2208] a speech recognition means for converting a voice message into text;
[2209] A means to store conversation history and user preferences in a database and load them as needed;
[2210] A system including:
[2211] (Claim 2)
[2212] 10. The system of claim 1, further comprising means for receiving user customization settings to adjust the tone of the dialogue and the reactions of the characters.
[2213] (Claim 3)
[2214] 10. The system of claim 1, further comprising means for performing initialization of the chat AI model and the speech generation model and acting based on user interaction.
[2215] "Application Example 1"
[2216] (Claim 1)
[2217] means for receiving input messages from a user;
[2218] A means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message;
[2219] means for initializing a voice generation model for converting the response message into voice data and generating voice data;
[2220] means for transmitting the generated audio data to a user;
[2221] A means for recreating a fictional character based on a user's selection within the virtual environment to provide guidance and product information;
[2222] A system including:
[2223] (Claim 2)
[2224] A means to verify the user's credentials and display the dashboard if authentication is successful;
[2225] A means for loading and initializing the data of a fictional character selected by the user;
[2226] A means for fictional characters to respond in real time when users request product details or purchasing advice within a virtual environment;
[2227] The system of claim 1 further comprising:
[2228] (Claim 3)
[2229] a speech recognition means for converting a user's voice message into a text message;
[2230] A means to store conversation history and user preferences in a database and load them as needed;
[2231] The system of claim 1 further comprising:
[2232] "Example 2: Combining Emotion Engines"
[2233] (Claim 1)
[2234] means for receiving input messages from a user;
[2235] means for initializing a dialogue generation model using a natural language processing algorithm for responding to the input message and generating a response message;
[2236] means for initializing a voice synthesis model for converting the response message into voice data and generating voice data;
[2237] means for utilizing a sentiment analysis engine to analyze a user's sentiment from the input message;
[2238] means for generating a response message taking into account the emotional status obtained by the emotion analysis engine;
[2239] a means for utilizing a speech synthesis model to generate speech data with a tone according to the emotional status;
[2240] means for transmitting the generated audio data to a user;
[2241] A system including:
[2242] (Claim 2)
[2243] A means to verify the user's credentials and display the dashboard if authentication is successful;
[2244] means for loading and initializing a dialogue generation model, a speech synthesis model, and a sentiment analysis engine corresponding to a fictional character selected by a user;
[2245] The system of claim 1 further comprising:
[2246] (Claim 3)
[2247] a speech recognition means for converting a user's voice message into a text message;
[2248] A means to store conversation history and user preferences in a database and load them as needed;
[2249] The system of claim 1 further comprising:
[2250] "Application example 2 when combining emotion engines"
[2251] (Claim 1)
[2252] means for receiving input messages from a user;
[2253] means for initializing a dialogue AI model using a natural language processing algorithm for responding to the input message and generating a response message;
[2254] means for initializing a voice generation algorithm for converting the response message into voice data and generating the voice data;
[2255] a means for adding emotion to a response message using an emotion analysis engine that analyzes the emotion of a user;
[2256] means for transmitting the generated audio data to a user;
[2257] A system including:
[2258] (Claim 2)
[2259] A means to verify the user's credentials and display the dashboard if authentication is successful;
[2260] A means for loading and initializing the data of a fictional character selected by the user;
[2261] A means of analyzing the user's emotional status and reflecting it in the initial settings;
[2262] The system of claim 1 further comprising:
[2263] (Claim 3)
[2264] a speech recognition means for converting a user's voice message into a text message;
[2265] A means to store conversation history and user preferences in a database and load them as needed;
[2266] A means to customize how fictional characters respond to user choices; and
[2267] The system of claim 1 further comprising: [Explanation of symbols]
[2268] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving input messages from a user; A means for initializing a chat AI model using a natural language processing algorithm to respond to the input message and generating a response message; means for initializing a voice generation model for converting the response message into voice data and generating voice data; means for transmitting the generated audio data to a user; A system including:
2. A means to verify the user's credentials and display the dashboard if authentication is successful; A means for loading and initializing the data of a fictional character selected by the user; The system of claim 1 further comprising:
3. a speech recognition means for converting a user's voice message into a text message; A means to store conversation history and user preferences in a database and load them as needed; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A