System
The system provides personalized and realistic conversations through generative AI, addressing loneliness and managing conversation data rights, enhancing user experience and reducing copyright risks.
Patent Information
- Application Number
- JP2024123938
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional dialogue systems fail to provide realistic and personalized conversations, exacerbating feelings of loneliness and isolation, and lack effective solutions for managing conversation data rights and reducing copyright infringement risks.
A system that allows users to select a specific person and engage in a realistic conversation through generative AI, which mimics that person's personality and speaking style, includes features for user authentication, data storage, and secondary use of interaction data.
Enables personalized and realistic conversations, reducing loneliness, and allows users to save and reuse dialogue data, thereby enhancing user experience and mitigating copyright risks.
Smart Images

Figure 2026022421000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, people are feeling increasingly isolated due to a declining birthrate, an aging population, the trend toward nuclear families, and urban living. In particular, an increasing number of people, such as mothers raising children and lonely elderly people, are seeking someone to talk to on a daily basis, but effective solutions to address this need are lacking. Furthermore, conventional dialogue systems struggle to provide a realistic conversational experience, limiting their effectiveness in alleviating loneliness and isolation. Furthermore, issues remain regarding the rights management of conversation data, necessitating appropriate reward distribution and reducing the risk of infringement. [Means for solving the problem]
[0005] The present invention provides a system that allows a user to select a specific person and experience a realistic conversation through a generative AI that imitates the person's personality and speaking style. The system includes the following means.
[0006] 1. A way for users to select a specific person and experience a realistic interaction through generative AI that mimics that person's personality and speaking style.
[0007] 2. A means of receiving text messages entered by the user and sending them to the server.
[0008] 3. A means by which the server generates an appropriate response based on the message it receives and sends it back to the terminal.
[0009] 4. A means for the server to store the interaction data between the user and the AI and provide it in a reusable format as needed.
[0010] 5. A means for the user to end the interaction and log out.
[0011] The system further includes means for the server to retrieve conversation data related to the selected person from the database, load and configure a corresponding AI model, and set the personality and speaking characteristics of the target person in the AI model to complete initialization.
[0012] The system also includes a means for the server to verify the user's authentication information, start a user session if the login is successful, and return an error message if the login fails. This provides a system that can personalize the experience of each individual user and effectively solve the problem of social isolation.
[0013] "User" refers to an individual who selects a specific person to interact with through the system and sends a message to the generating AI.
[0014] "Specific person" refers to the person the generative AI will imitate, such as a celebrity, public figure, family member, or fictional character with whom the user wishes to interact.
[0015] "Personality" refers to a person's unique character traits, characteristics, and speaking style.
[0016] "Generative AI" refers to a program that uses artificial intelligence technology to mimic the personality and speaking style of a specific person and generate realistic dialogue.
[0017] "Server" refers to the central computer system that receives user messages, generates responses using generative AI, and manages and stores dialogue data.
[0018] "Terminal" means the device (e.g., computer, smartphone, tablet) through which a User accesses the System and sends and receives Messages.
[0019] A "message" refers to text information that a user enters and sends.
[0020] "Authentication information" refers to information such as a username and password provided when a user logs in.
[0021] A "session" refers to a series of interactions and operations while a user is logged into a system.
[0022] "Interaction data" refers to a record of a series of messages and responses between a user and the generating AI.
[0023] "Secondary use" refers to using saved dialogue data as audio files or other applications.
[0024] "Logging out" refers to the action of a user exiting the system and ending the session. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0026] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0027] First, the terms used in the following description will be explained.
[0028] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0029] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0030] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0031] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0033] [First embodiment]
[0034] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0035] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0036] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0037] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0038] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0040] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0041] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0042] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0043] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0044] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0045] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0046] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[0047] Overall system overview
[0048] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. In addition, the conversation data from the system is saved and can be reused, reducing the risk of copyright infringement.
[0049] System Components and Operation
[0050] 1. User Device
[0051] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0052] 2. Server
[0053] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[0054] Program processing flow
[0055] System startup and user authentication
[0056] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0057] Character selection and AI settings
[0058] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0059] Start a dialogue
[0060] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[0061] Response display and conversation continuation
[0062] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0063] Storage and secondary use of conversation data
[0064] The server stores the conversation data between the user and the generated AI in a database in real time as a log. Users can reuse their conversation data as needed, such as creating audio files, and the server converts the data for this purpose.
[0065] Exit and log out
[0066] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0067] Specific examples
[0068] For example, consider the process of a user logging in, selecting a particular famous author, and starting a dialogue. When the user types and sends a question such as, "What do you think of the latest novel?", the server passes the message to an AI model, which generates an appropriate response (e.g., "I've tried a particular subject. I'm curious to know how it resonates with you.") and sends it back to the device. The device displays this response to the user, and the dialogue continues. This series of processes allows the user to have a realistic dialogue experience.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, user interface module) and connects to the database.
[0072] Step 2:
[0073] The user enters a username and password into the login page on the terminal.
[0074] Step 3:
[0075] The terminal transmits the login information entered by the user to the server.
[0076] Step 4:
[0077] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is taken to the main screen.
[0078] Step 5:
[0079] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[0080] Step 6:
[0081] The terminal transmits the information of the selected person to the server.
[0082] Step 7:
[0083] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking style to complete the initialization.
[0084] Step 8:
[0085] The user enters a text message and presses the send button to initiate a conversation.
[0086] Step 9:
[0087] The terminal transmits the input message to the server.
[0088] Step 10:
[0089] The server passes the received message to the AI model, which generates the optimal response to that message.
[0090] Step 11:
[0091] The server sends the generated response message to the terminal.
[0092] Step 12:
[0093] The terminal displays the response message received from the server to the user.
[0094] Step 13:
[0095] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0096] Step 14:
[0097] The server stores the interaction data between the user and the generated AI in a database in real time.
[0098] Step 15:
[0099] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[0100] Step 16:
[0101] The terminal sends a request to end the conversation to the server.
[0102] Step 17:
[0103] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[0104] Example 1
[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0106] In recent years, dialogue systems using generative AI that mimic the personality and speaking style of a specific person have been attracting attention, but existing systems have not been able to provide realistic and personalized dialogue to reduce users' feelings of loneliness, limiting the experience.In addition, functions related to saving and secondary use of dialogue data are limited, presenting technical challenges for further enriching the user experience.
[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0108] In this invention, the server includes means for initializing modules and connecting to a database when the system starts up, means for verifying user authentication information against the database and starting a user session if authentication is successful, means for allowing a user to select a specific person, acquiring conversation data related to that person, and loading and configuring a corresponding AI model, and means for storing the dialogue data in the database in real time and providing it in a reusable format as needed, thereby enabling users to smoothly engage in real, personalized dialogue with specific people and facilitating the storage and secondary use of the dialogue data.
[0109] A "user" refers to an entity that uses a specific terminal to access the dialogue system and engage in dialogue.
[0110] A "terminal" is a device used by a user to interact, and includes a PC, smartphone, tablet, etc.
[0111] A "server" is a computer system that plays a central role in the system, performing user authentication, managing interaction data, and operating the generation AI.
[0112] "Generative AI" refers to artificial intelligence models that mimic the personality and speaking style of a specific person to provide realistic dialogue.
[0113] "Personality" refers to a person's unique character traits and distinctive style of expression.
[0114] "Speech style" refers to a particular person's tone of voice and the way they use language.
[0115] A "prompt sentence" refers to a text message entered into the generation AI, based on which the AI generates a response.
[0116] A "user session" refers to session information for managing a series of interactions while a user is logged into a system.
[0117] A "database" is a system for storing and managing various data such as dialogue data and user information.
[0118] A "module" refers to a program component for realizing a specific function of a system.
[0119] "Initialization" refers to the process by which a system or module prepares to start operation.
[0120] "Logging out" refers to a user exiting the system and ending the user session.
[0121] "Interaction Data" refers to records of messages exchanged between the user and the generating AI.
[0122] "Secondary use" refers to converting the saved dialogue data into another format, such as an audio file, or using it for other purposes.
[0123] This section describes a specific embodiment of the present invention. This system allows users to select a specific person and provides realistic dialogue through a generative AI that mimics that person's personality and speaking style. Each of the following steps details the configuration and operation of the respective hardware and software.
[0124] System Configuration
[0125] The system mainly consists of the following components:
[0126] 1. User Device
[0127] A terminal is a device with which a user interacts, and includes a PC, smartphone, tablet, etc. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0128] 2. Server
[0129] The server plays a central role in the system, performing user authentication, dialogue data management, and operation management of the generative AI. The server is composed of multiple modules (e.g., a user authentication module, a data management module, an AI engine, etc.).
[0130] System Operation
[0131] The main system operations are explained below.
[0132] 1. System startup and user authentication
[0133] When the server starts the system, all necessary modules are initialized and connected to the database.
[0134] When a user enters login information on a terminal, the terminal sends that information to the server, which checks the received authentication information against a database and starts a user session if authentication is successful. If authentication fails, it returns an error message to the terminal.
[0135] 2. Character selection and AI settings
[0136] The user selects a specific person with whom they wish to interact on the platform, and the device sends the selected person's information to the server.
[0137] Based on the received person ID, the server retrieves the conversation data related to that person from the database, loads the corresponding AI model, and sets the selected person's personality and speaking style to the AI model, completing the initialization.
[0138] 3. Start the dialogue
[0139] The user inputs a text message to start a conversation and presses the send button. The terminal sends the input message to the server.
[0140] The server passes the received message to the AI model, which generates the optimal response to that message, which is then sent back to the device, where it is displayed to the user.
[0141] 4. Response display and conversation continuation
[0142] When the user enters and sends a new message to continue the conversation, the terminal again sends the message to the server.
[0143] The server again passes the message to the AI model and generates a response, and this process repeats until the interaction is over.
[0144] 5. Storage and secondary use of conversation data
[0145] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[0146] Users can reuse their dialogue data as audio files, etc., as needed, and the server converts the data accordingly. For example, to convert text data into audio data, a Text-to-Speech (TTS) engine is used.
[0147] 6. Exiting and Logging Out
[0148] When the user wants to end the conversation, he clicks the end button, and the terminal sends a request to end the conversation to the server.
[0149] The server ends the session, stores any necessary logs, and finally logs the user out and exits the system.
[0150] Specific examples
[0151] For example, consider the process of a user logging in and selecting a particular famous person (e.g., a famous scientist) to begin a dialogue. When the user types and submits a question such as, "What do you think about the latest scientific research?", the server passes the message to an AI model, which generates an appropriate response (e.g., "The latest research is very interesting, especially its original approach.") and sends it back to the device. The device then displays this response to the user and continues the dialogue.
[0152] Prompt Sentence Examples
[0153] For example, you might input the following prompts into a generative AI model:
[0154] "What do you think about the latest scientific research?"
[0155] The prompt allows the user to ask a question to the AI model as a specific person, with the expectation that it will generate a response that mimics the way that person speaks and thinks.
[0156] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0157] Step 1: System startup and module initialization
[0158] When the server starts the system, it initializes all necessary modules (user authentication module, data management module, AI engine, etc.) and connects to the database.
[0159] Input: System startup trigger.
[0160] Processing: Module initialization, database connection establishment.
[0161] Output: System initialization complete and database connection status.
[0162] Step 2: User authentication
[0163] The user enters login information on the device and clicks the login button.
[0164] The terminal sends the entered login information (user name, password) to the server.
[0165] The server checks the received authentication information against the user information in its database, and if authentication is successful, starts a user session.
[0166] Input: Username, Password.
[0167] Processing: Database verification, session start.
[0168] Output: Session ID, authentication success or failure result.
[0169] Step 3: Selecting a person and setting AI
[0170] The user selects the specific person with whom they wish to interact on the platform.
[0171] The device sends the selected person's information (person ID and name) to the server.
[0172] Based on the received person information, the server retrieves conversation data related to that person from the database and loads the corresponding AI model.
[0173] The server sets the AI model with the selected person's personality and speaking style, completing initialization.
[0174] Input: Person ID, Name.
[0175] Processing: Acquire conversation data, load AI model, and set personality.
[0176] Output: The initialized AI model.
[0177] Step 4: Start a conversation
[0178] The user enters a text message to initiate a conversation and presses the send button.
[0179] The terminal transmits the input message to the server.
[0180] The server passes the received message to the AI model, which generates the optimal response to that message.
[0181] The generated response is sent back from the server to the terminal, which displays the response to the user.
[0182] Input: User message.
[0183] Processing: Parse the message, feed it into the AI model, and generate a response.
[0184] Output: The generated response message.
[0185] Step 5: Keep the conversation going
[0186] The user enters and sends a new message to continue the conversation.
[0187] The terminal sends a new message to the server.
[0188] The server again passes the message to the AI model, which generates a response.
[0189] This process is repeated until the conversation is over.
[0190] Input: New user message.
[0191] Processing: Parse the message, feed it into the AI model, and generate a response.
[0192] Output: The generated response message.
[0193] Step 6: Saving and Reusing Conversation Data
[0194] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[0195] The user can reuse the dialogue data as a voice file or the like.
[0196] The server performs the data conversion for this purpose, for example by using a Text-to-Speech (TTS) engine.
[0197] Input: Interaction data.
[0198] Processing: Data storage, data conversion (e.g. text to speech).
[0199] Output: Saved interaction logs, transformed data.
[0200] Step 7: End the conversation and log out
[0201] If the user wishes to end the interaction, he clicks on the End button.
[0202] The terminal sends a request to end the conversation to the server.
[0203] The server terminates the session and saves any necessary logs.
[0204] Finally, the server logs the user out and causes him to leave the system.
[0205] Input: A conversation termination request.
[0206] Processing: Session termination, log saving, logout processing.
[0207] Output: Finished sessions, saved interaction logs, logout status.
[0208] (Application example 1)
[0209] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0210] Conventional dialogue systems allow users to have a realistic dialogue experience through dialogue with a generative AI that mimics the personality and speaking style of a specific person. However, there is a need to provide a richer dialogue experience by adding entertainment and information functions. Furthermore, there is a lack of functionality that allows users to save their dialogue history and use it for later playback or learning. Technology is needed to solve these issues.
[0211] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0212] In this invention, the server includes means for retrieving conversation data related to a selected person from a database and loading and configuring a corresponding generative AI model, means for verifying user authentication information and starting a user session if login is successful, and means for adjusting and optimizing the generative AI model for providing entertainment and information. This allows a user to enjoy entertainment and information in real time through conversations with a specific person, and saves the conversation history for later playback or learning.
[0213] A "user" is an entity that uses the system to have a conversation with a specific person.
[0214] A "specific person" is someone the user selects and a dialogue is provided through a generative AI that mimics that person's personality and speaking style.
[0215] "Personality" refers to the characteristics, traits, and speaking style that are unique to a particular person.
[0216] "Generative AI" is artificial intelligence that mimics the personality and speaking style of a specific person and generates realistic conversations with the user.
[0217] A "text message" is the text information entered by the user, and is the data that the generation AI uses to generate a response.
[0218] A "server" is a device that plays a central role in the system, performing user authentication, managing dialogue data, and operating the generation AI.
[0219] A "response" is a reply message that the generation AI generates in response to a user's text message.
[0220] A "terminal" is a device used by a user to interact, and includes a smartphone, a personal computer, etc.
[0221] "Dialogue data" is a record of the conversation between the user and the generating AI.
[0222] "Secondary use" refers to reusing saved dialogue data in a different format or for a different purpose.
[0223] "Logging out" means ending an interactive session and exiting the system by the user.
[0224] "Conversation data" is dialogue information related to a particular person.
[0225] A "database" is a system that stores and manages information such as conversation data.
[0226] An "AI model" is an algorithm or mathematical model that serves as the basis for generative AI that mimics the personality and speaking style of a specific person.
[0227] "Entertainment" refers to interactive content and functions that provide enjoyment to users.
[0228] "Information" refers to knowledge and data that users acquire through dialogue.
[0229] "Playback" refers to using the saved dialogue history to check it later.
[0230] "Logging in" is the process by which a user enters and confirms authentication information in order to access and use a system.
[0231] "Authentication information" refers to personal information required to access a system, such as a user ID and password.
[0232] A "user session" is a series of operations performed by a user from the time the user logs in until the time the user logs out.
[0233] An "error message" is a warning or guidance message that is displayed to the user when login fails.
[0234] "Tuning" is the process of configuring and optimizing a generative AI model.
[0235] "Optimization" refers to adjusting the generative AI model to improve the user's interaction experience.
[0236] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[0237] Overall system overview
[0238] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly useful for parents raising children and lonely elderly people. In addition, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement.
[0239] System Components and Operation
[0240] 1. User Device
[0241] The terminal is a device with which the user interacts, and can be a smartphone, smart glasses, or a head-mounted display. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0242] 2. Server
[0243] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[0244] Program processing flow
[0245] 1. System startup and user authentication
[0246] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0247] 2. Character selection and AI settings
[0248] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0249] 3. Start the dialogue
[0250] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[0251] 4. Response display and conversation continuation
[0252] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0253] 5. Storage and secondary use of conversation data
[0254] The server stores the dialogue data between the user and the generated AI in a database in real time as a log. The user can reuse the dialogue data as audio files, etc., as needed, and the server converts the data for this purpose.
[0255] 6. Exiting and Logging Out
[0256] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0257] Hardware and software used
[0258] Hardware: Smartphones, smart glasses, head-mounted displays, servers
[0259] Software: Web applications, native applications, generative AI models, databases (MySQL, PostgreSQL), cloud services (AWS, Google Cloud)
[0260] Examples of specific examples and prompts
[0261] For example, consider a scenario in which a user logs into a "virtual conversation hall" and begins a conversation with a famous author. When the user types and sends a question such as "What do you think about the latest novel?", the server passes the message to an AI model, which generates an appropriate response. The generated response (e.g., "I've tried a particular theme. I'm curious to see how it resonates with you.") is sent back to the device and displayed to the user. In this way, the user can have a realistic conversation experience.
[0262] Prompt Sentence Examples
[0263] User: What do you think about the latest novel?
[0264] Server: Sends messages to the AI model and generates responses such as:
[0265] AI Model: "I've tried a particular theme and I'm curious to see how it resonates."
[0266] In this way, the present invention provides users with a realistic interaction experience with a specific person, enabling them to enjoy entertainment and information.
[0267] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0268] Step 1:
[0269] System startup and user authentication:
[0270] The server starts the system, initializes all necessary modules, and connects to the database. The user enters login information (username, password) on the terminal. The terminal sends the entered information to the server. The server compares the received login information with the database, and if authentication is successful, starts the user session. At this time, it obtains the user's authentication information from the database after hashing it and compares it with the entered information. If authentication fails, an error message is returned to the user's terminal. As output, if successful, the user session will start, and if unsuccessful, an error message will be displayed.
[0271] Step 2:
[0272] Character Selection and AI Settings:
[0273] The user selects a specific person with whom they wish to interact on the platform. The device sends the selected person's information to the server. The server uses this information to retrieve conversation data related to the selected person from a database. The server then loads the corresponding generative AI model, sets the selected person's personality and speaking style in the AI model, and completes initialization. The configured generative AI model is prepared as the output.
[0274] Step 3:
[0275] Start the conversation:
[0276] The user enters a text message to start a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to a generative AI model, which generates an optimal response to that message. In this case, the generative AI model (e.g., OpenAI GPT-3, GPT-4) takes text data as input and runs an algorithm to generate an appropriate response. The generated response is sent back from the server to the device, and the device displays it to the user. As an output, a response message is generated that is displayed on the user's device.
[0277] Step 4:
[0278] Response display and conversation continuation:
[0279] The user types and sends a new text message to continue the dialogue. The device again sends the message to the server. The server again passes the message to the generative AI model, which generates a new response. This process repeats until the dialogue ends. Each message exchange involves data processing and the generative AI generating a response based on the new user input. As an output, the continuing dialogue message is displayed on the user's device.
[0280] Step 5:
[0281] Conversation data storage and secondary use:
[0282] The server stores the dialogue data between the user and the generated AI in a database in real time. The stored dialogue data can be converted into an audio file or other format as needed for secondary use. The server uses a data conversion algorithm to convert text data into an audio file. As output, the stored dialogue data and data in a format that can be used for secondary use are generated.
[0283] Step 6:
[0284] Exit and log out:
[0285] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server. The server ends the session and saves the necessary logs. Finally, the user is logged out and exits the system. As an output, the user is safely logged out of the system.
[0286] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0287] As an embodiment of the present invention, a system is described below that combines an emotion engine with a dialogue platform using generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue. This system is realized when a user selects a specific person and sends a message to the generative AI, which recognizes the user's emotions and generates and provides a response based on them.
[0288] Overall system overview
[0289] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through generative AI. This system is designed to help reduce the user's sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. Furthermore, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly.
[0290] System Components and Operation
[0291] 1. User Device
[0292] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0293] 2. Server
[0294] The server plays a central role in the system, performing user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[0295] 3. Generation AI
[0296] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue. The server sets the necessary parameters and generates a response based on the user's message.
[0297] 4. Emotion Engine
[0298] The emotion engine recognizes emotions from the text messages entered by the user and adjusts the generative AI's response based on those emotions. Emotion data is also stored along with the dialogue data and can be used for analysis and feedback as needed.
[0299] Program processing flow
[0300] System startup and user authentication
[0301] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0302] Character selection and AI settings
[0303] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0304] Start a dialogue
[0305] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to an emotion engine to recognize the user's emotion. The recognized emotion data is passed to an AI model, which generates a tailored response based on the emotion. The generated response is sent back to the device, which then displays it to the user.
[0306] Response display and conversation continuation
[0307] When the user enters and sends a new message to continue the conversation, the device again sends the message to the server, which then passes it to the emotion engine, which recognizes the user's emotion and generates an appropriate response. This process is repeated until the conversation ends.
[0308] Storage and secondary use of conversation data
[0309] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time. Users can reuse their conversation data as audio files, etc., as needed, and the server converts the data for this purpose. The server also analyzes the emotional data and uses it as feedback to improve the quality of the conversation.
[0310] Exit and log out
[0311] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0312] Specific examples
[0313] For example, when a user logs in and selects a particular famous author to begin a dialogue, the user can type in a question like, "What do you think about the latest novel?" and send it. The server then passes the message to the emotion engine, which recognizes the user's emotion (e.g., excitement, curiosity). Based on the emotion data, the AI model generates a response (e.g., "I've tried a particular topic. I'm curious to see how it resonates with you.") and sends it back to the device. The device then displays this response to the user, continuing the dialogue. This process allows the user to have a more personal and realistic dialogue experience.
[0314] The processing flow will be explained below.
[0315] Step 1:
[0316] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, emotion engine module, and user interface module) and connects to the database.
[0317] Step 2:
[0318] The user enters a username and password into the login page on the terminal.
[0319] Step 3:
[0320] The terminal transmits the login information entered by the user to the server.
[0321] Step 4:
[0322] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is redirected to the main screen. If authentication fails, an error message is returned.
[0323] Step 5:
[0324] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[0325] Step 6:
[0326] The terminal transmits the information of the selected person to the server.
[0327] Step 7:
[0328] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking characteristics to complete initialization.
[0329] Step 8:
[0330] The user enters a text message and presses the send button to initiate a conversation.
[0331] Step 9:
[0332] The terminal transmits the input message to the server.
[0333] Step 10:
[0334] The server passes the received message to the emotion engine, which analyzes the text and recognizes the user's emotion from the message.
[0335] Step 11:
[0336] The server passes the recognized user emotion data to the AI model, which uses it to generate an appropriate response.
[0337] Step 12:
[0338] The server generates a response message adjusted based on the emotion data and transmits it to the terminal.
[0339] Step 13:
[0340] The terminal displays the response message received from the server to the user.
[0341] Step 14:
[0342] When the user enters and sends a new message to continue the conversation, the terminal sends the message to the server again.
[0343] Step 15:
[0344] The server again passes the message to the emotion engine, which recognizes the emotion and generates a response based on the recognized emotion data. This process is repeated until the interaction is completed.
[0345] Step 16:
[0346] The server stores the conversation data and emotion data between the user and the generated AI in a database in real time.
[0347] Step 17:
[0348] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[0349] Step 18:
[0350] The terminal sends a request to end the conversation to the server.
[0351] Step 19:
[0352] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[0353] Example 2
[0354] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0355] In modern society, many users experience feelings of loneliness, a major problem especially for mothers raising children and lonely elderly people. Conventional dialogue systems struggle to fully recognize users' emotions and respond appropriately, preventing users from engaging in satisfying dialogue. Furthermore, data integrity and security are not guaranteed when dialogue data is stored or reused. A system that can address these issues is needed.
[0356] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotions using an emotion engine based on a message received by the server and generating an appropriate response, means for returning the generated response to the terminal, and means for the server to store dialogue data and emotion data between the user and the generation AI and provide them in a format that can be reused. This allows the user to experience a realistic dialogue and obtain an appropriate response based on the emotion. Furthermore, since the stored dialogue data and emotion data can be reused, it is possible to improve the quality of the dialogue and increase user satisfaction.
[0357] "User" refers to a person who uses the system to interact with a specific person.
[0358] "Terminal" refers to a device such as a PC, smartphone, or tablet that a user uses to interact.
[0359] A "server" refers to a computer that plays a central role in the system and performs user authentication, dialogue data management, generative AI operation, emotion recognition, etc.
[0360] "Generative AI" refers to an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[0361] An "emotion engine" refers to a program that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[0362] "Dialogue data" refers to data that records the content of the dialogue between the user and the generating AI.
[0363] "Secondary use" refers to converting saved dialogue data and emotion data into audio files, etc., and reusing them for analysis or other purposes.
[0364] "Initialization" refers to the process of setting up a system or program so that it can operate normally.
[0365] "Personality" refers to the characteristics and traits that a particular person possesses.
[0366] "Logging in" refers to the process by which a user enters authentication information to access a system and is authenticated by the system.
[0367] A "session" refers to a series of interactions with the system from the time a user logs in until the time they log out.
[0368] This invention is a dialogue platform system that uses generative AI to mimic the personality and speaking style of a specific person and provide realistic dialogue. When a user selects a specific person and sends a message to the generative AI, the AI recognizes the user's emotions and generates and provides a response based on those emotions. This system is particularly useful for users who feel lonely, such as mothers raising children or lonely elderly people.
[0369] System Components
[0370] 1. User Device
[0371] Terminals include devices connected to the Internet, such as PCs, smartphones, tablets, etc. Users interact with these terminals.
[0372] The terminal accesses the server through a web application or a native application.
[0373] 2. Server
[0374] The server is the core of the system, and is responsible for user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[0375] The server sets the necessary parameters and generates a response according to the user's message.
[0376] 3. Generation AI
[0377] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[0378] The generative AI creates a realistic interaction experience for users based on the characteristics of the person they select.
[0379] 4. Emotion Engine
[0380] The emotion engine is responsible for recognizing emotions from the text messages entered by the user and adjusting the generative AI's response based on those emotions.
[0381] Emotional data is stored along with the dialogue data and can be used for analysis and feedback as needed.
[0382] System Operation
[0383] In this system, the user initiates a conversation using a device, and the server provides a realistic conversation via a generative AI. The overall flow of the system is shown below.
[0384] 1. User authentication: The user enters login information on the terminal, which then sends it to the server. The server checks the authentication information against a database and starts the session if authentication is successful. If authentication fails, it returns an error message.
[0385] 2. Person selection: The user selects the desired conversation partner, and the device sends this information to the server. The server retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The AI model initialization is complete.
[0386] 3. Interaction: When the user inputs a message, the device sends it to the server. The server passes the message to the emotion engine to recognize the user's emotion. It generates a response tailored based on the emotion and sends it back to the device. The device displays the response to the user.
[0387] 4. Continuous Dialogue: This process repeats as long as the user continues to input new messages. For each message, the appropriate emotion is recognized and a response is generated.
[0388] 5. Data storage and secondary use: The server stores the conversation data and emotion data between the user and the generated AI in a database. The user can reuse the conversation data as audio files, and the server converts the data for this purpose. The server also analyzes the emotion data and uses it as feedback to improve the quality of the conversation.
[0389] 6. Termination and logout: When the user finishes the conversation, the terminal sends a request to terminate the conversation to the server, which then terminates the session, saves the necessary logs, and logs out the user.
[0390] Specific examples of operation
[0391] For example, a user logs into the system, selects a particular well-known author, and begins a dialogue by typing and sending a question such as, "What do you think of his latest novel?" The server passes this message to the emotion engine, which recognizes the user's emotion (e.g., excitement or curiosity). Based on this emotion data, the generative AI generates a response such as, "I've tried a particular theme. I'm curious to see how it resonates with you," and sends it back to the device. The device displays this response to the user, allowing the user to continue the dialogue.
[0392] Prompt Sentence Examples
[0393] "What is your opinion on the latest science and technology?"
[0394] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0395] Step 1:
[0396] System startup
[0397] The server initializes all modules in the system, securing the necessary resources and establishing connections to the database.
[0398] Input: System startup instructions
[0399] Data processing: Module initialization, resource allocation
[0400] Output: Initialization complete, connection established
[0401] Step 2:
[0402] User Authentication
[0403] The user enters login information (user name, password) on the terminal and presses the send button.
[0404] The terminal transmits the entered authentication information to the server.
[0405] The server receives the input information and checks it against a database.
[0406] Input: Authentication information (username, password)
[0407] Data processing: Matching with database
[0408] Output: Authentication success or failure message
[0409] If authentication is successful, the server starts a user session and generates a corresponding session ID, otherwise it returns an error message to the terminal.
[0410] Step 3:
[0411] Character Selection
[0412] Users select the specific person on the platform they want to interact with.
[0413] The terminal transmits the information of the selected person to the server.
[0414] Based on the received information, the server retrieves relevant conversation data from the database and loads the corresponding AI model.
[0415] Input: Selected person's information
[0416] Data processing: Acquire conversation data, load AI model
[0417] Output: AI model initialization complete
[0418] Step 4:
[0419] Start a dialogue
[0420] The user enters a message and presses the send button.
[0421] The terminal transmits the input message to the server.
[0422] The server passes the received message to the emotion engine to recognize the user's emotion.
[0423] Input: User message
[0424] Data processing: Emotion recognition
[0425] Output: Emotion data
[0426] The generation AI generates a response based on the recognized emotion data, and the server sends the generated response back to the device.
[0427] Step 5:
[0428] Response Display
[0429] The terminal displays the response received from the server to the user.
[0430] Input: Generator AI response
[0431] Data processing: Displaying responses
[0432] Output: The response displayed on the user's screen
[0433] Step 6:
[0434] Ongoing dialogue
[0435] The user enters a new message and repeats steps 4 and 5 above.
[0436] Input: New user message
[0437] Data processing: emotion recognition, response generation
[0438] Output: Shows the new response
[0439] Step 7:
[0440] Saving conversation data
[0441] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time.
[0442] Input: Dialogue data, emotion data
[0443] Data processing: Saving to database
[0444] Output: Saved interaction data
[0445] Step 8:
[0446] Secondary use
[0447] The user can reuse the saved dialogue data as a voice file or the like.
[0448] The server converts the interaction data and provides it in the specified format.
[0449] Input: Saved interaction data
[0450] Data processing: Data conversion
[0451] Output: Converted data format
[0452] Step 9:
[0453] Exit and log out
[0454] If the user wishes to end the dialogue, he or she presses the end button.
[0455] The terminal sends a request to end the conversation to the server.
[0456] The server ends the session, saves any necessary logs, and logs the user out.
[0457] Input: Conversation termination request
[0458] Data processing: End session, save log
[0459] Output: Logout completed
[0460] (Application example 2)
[0461] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0462] Current dialogue systems for factory robots have difficulty achieving natural communication with humans and are unable to provide sufficient support to workers. As a result, it is difficult to improve work efficiency and safety, and there is a risk of worker fatigue and stress increasing. In addition, the robot lacks the ability to recognize workers' emotions and respond appropriately based on them, which means that there is an issue of not being able to reduce workers' psychological burden or provide sufficient support for their work.
[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotion and generating a response based on that emotion, means for the server to recognize the user's emotion and adjust the response based on the emotion data, and means for the user to receive work instructions through dialogue with the robot. This not only enables natural communication when factory workers interact with the robot, but also enables appropriate responses and support that take into account the worker's emotion and fatigue. As a result, it is expected that the psychological burden on workers will be reduced and work efficiency and safety will be improved.
[0464] "User" refers to a human being who uses the system to engage in dialogue using generative AI.
[0465] A "specific person" refers to a person selected by the user whose personality and speaking style the generated AI will imitate and set up to converse with.
[0466] "Generative AI" refers to artificial intelligence programs that are designed to mimic the personality and speaking style of a specific person.
[0467] "Server" refers to the central computer system that manages the storage of interaction data, user authentication, emotion recognition, and the operation of generative AI.
[0468] An "emotion engine" refers to a module that recognizes emotions from messages entered by users and generates responses based on those emotions.
[0469] A "terminal" is a device with which a user interacts, such as a smartphone, tablet, or computer.
[0470] "Dialogue data" refers to information including messages and emotional data exchanged between the user and the generating AI.
[0471] "Secondary use" refers to reusing saved dialogue data for other purposes.
[0472] "AI model" refers to an artificial intelligence module that includes the configuration information and algorithms used by the generative AI.
[0473] "Emotion data" refers to data that recognizes a user's emotions and includes information about them.
[0474] A "robot" refers to a mechanical device used to interact with workers in factories and other places.
[0475] "Work instructions" refer to instructions regarding specific work content that a user receives through interaction with a robot.
[0476] "Login information" refers to information used for user authentication, and primarily refers to user IDs and passwords.
[0477] This invention is a system that uses a support robot to interact with factory workers. The system is designed to allow users to experience realistic interactions through a generative AI that mimics the personality and speaking style of a specific person. The specific configuration and operation for implementing this system are described below.
[0478] Overall system configuration
[0479] The system mainly consists of the following components:
[0480] 1. User Device:
[0481] The devices used are smartphones, tablets, and computers. These devices connect to the internet and communicate with the server.
[0482] 2. Server:
[0483] The Server is the central computer system that processes user input messages and generates responses using generative AI. The Server contains the following main modules:
[0484] User Authentication Module
[0485] Generative AI Module
[0486] Emotion Engine Module
[0487] Database Management Module
[0488] 3. Generation AI:
[0489] Generative AI is an artificial intelligence program that generates realistic dialogue by imitating the personality and speaking style of a specific person, and is primarily implemented using the GPT-2 model.
[0490] 4. Emotion Engine:
[0491] The emotion engine is a module that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[0492] Operation overview
[0493] The system operation is outlined below.
[0494] 1. User authentication:
[0495] The user enters login information on the terminal and sends it to the server, which verifies the information and, if authentication is successful, starts the user session.
[0496] 2. Selecting a specific person:
[0497] The user selects a specific person with whom they wish to interact, and the device transmits the selection to the server, which loads and configures the corresponding generative AI model.
[0498] 3. Start a conversation:
[0499] To initiate a dialogue, a user inputs and sends a text message, and the server passes the message to an emotion engine, which recognizes the user's emotion, generates an appropriate response, and sends it back to the device.
[0500] 4. Continuing the conversation and storing data:
[0501] As the conversation continues, new messages from the user are processed in the same way. All conversation data is stored on the server in real time and made available in a reusable format.
[0502] Usage example
[0503] For example, imagine a factory worker logs into the system and selects a specific support representative to begin a conversation. If the user types, "Please tell me how to operate this machine," the following prompt sentence is generated:
[0504] Example prompt sentence:
[0505] The following is a conversation with a supportive robot. The robot is very helpful and supportive.
[0506] Human: Please tell me how to operate this machine.
[0507] Robot (supportive):
[0508] The server uses a generative AI model to generate a response, such as, "First, turn on the machine, then follow the instructions on the control panel."
[0509] The system allows factory workers to receive specific work instructions through a realistic interactive experience, which is expected to improve efficiency and safety.
[0510] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0511] Step 1:
[0512] User authentication: The user enters login information (user ID and password) on the terminal and sends it to the server. The server checks the information against a database, and if authentication is successful, a user session begins. If authentication is successful, the server generates a session ID and returns it to the terminal. Conversely, if authentication fails, the server returns an error message.
[0513] Input: User ID, Password
[0514] Data processing: Database matching of authentication information
[0515] Output: Session ID (if successful), Error message (if unsuccessful)
[0516] Step 2:
[0517] Selecting a specific person: The user selects a specific person with whom they wish to converse. The device sends the selected person's information to the server. The server retrieves the conversation data related to that person from the database, loads and configures the generative AI model, and then configures the model with the specific person's personality and speaking style.
[0518] Input: Selected person's information
[0519] Data processing: Acquiring a database of conversation data, loading and configuring the generative AI model
[0520] Output: The configured generative AI model
[0521] Step 3:
[0522] Dialogue start: The user enters a text message to start the dialogue and sends it from the device to the server. The server passes the received message to the emotion engine, which recognizes the user's emotion. The recognized emotion data is passed to a generative AI model, which generates an appropriate response based on the emotion. The generated response is then sent back to the device and displayed to the user.
[0523] Input: User's text message
[0524] Data processing: Emotion recognition and response generation using generative AI
[0525] Output: Response message
[0526] Step 4:
[0527] Continuing the dialogue and saving data: Every time the user inputs a new message to continue the dialogue, the device sends the message to the server. Emotion recognition is performed in the same way, and a response is generated based on the emotional data. This series of processes is then repeated. Furthermore, the server saves all dialogue data in a database.
[0528] Input: The user's new text message
[0529] Data processing: Emotion recognition, response generation using generative AI, and dialogue data storage
[0530] Output: Response message, saved interaction data
[0531] Step 5:
[0532] Ending the conversation and logging out: When the user wants to end the conversation, he clicks the end button and the terminal sends a request to end the conversation to the server. The server ends the conversation session and saves the log information. After that, the user is logged out and exits the system.
[0533] Input: Conversation termination request
[0534] Data processing: End session, save log information
[0535] Output: Logged out of the system
[0536] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0537] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0538] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0539] [Second embodiment]
[0540] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0541] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0542] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0543] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0544] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0545] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0546] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0547] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0548] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0549] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0550] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0551] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0552] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[0553] Overall system overview
[0554] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. In addition, the conversation data from the system is saved and can be reused, reducing the risk of copyright infringement.
[0555] System Components and Operation
[0556] 1. User Device
[0557] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0558] 2. Server
[0559] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[0560] Program processing flow
[0561] System startup and user authentication
[0562] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0563] Character selection and AI settings
[0564] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0565] Start a dialogue
[0566] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[0567] Response display and conversation continuation
[0568] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0569] Storage and secondary use of conversation data
[0570] The server stores the conversation data between the user and the generated AI in a database in real time as a log. Users can reuse their conversation data as needed, such as creating audio files, and the server converts the data for this purpose.
[0571] Exit and log out
[0572] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0573] Specific examples
[0574] For example, consider the process of a user logging in, selecting a particular famous author, and starting a dialogue. When the user types and sends a question such as, "What do you think of the latest novel?", the server passes the message to an AI model, which generates an appropriate response (e.g., "I've tried a particular subject. I'm curious to know how it resonates with you.") and sends it back to the device. The device displays this response to the user, and the dialogue continues. This series of processes allows the user to have a realistic dialogue experience.
[0575] The processing flow will be explained below.
[0576] Step 1:
[0577] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, user interface module) and connects to the database.
[0578] Step 2:
[0579] The user enters a username and password into the login page on the terminal.
[0580] Step 3:
[0581] The terminal transmits the login information entered by the user to the server.
[0582] Step 4:
[0583] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is taken to the main screen.
[0584] Step 5:
[0585] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[0586] Step 6:
[0587] The terminal transmits the information of the selected person to the server.
[0588] Step 7:
[0589] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking style to complete the initialization.
[0590] Step 8:
[0591] The user enters a text message and presses the send button to initiate a conversation.
[0592] Step 9:
[0593] The terminal transmits the input message to the server.
[0594] Step 10:
[0595] The server passes the received message to the AI model, which generates the optimal response to that message.
[0596] Step 11:
[0597] The server sends the generated response message to the terminal.
[0598] Step 12:
[0599] The terminal displays the response message received from the server to the user.
[0600] Step 13:
[0601] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0602] Step 14:
[0603] The server stores the interaction data between the user and the generated AI in a database in real time.
[0604] Step 15:
[0605] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[0606] Step 16:
[0607] The terminal sends a request to end the conversation to the server.
[0608] Step 17:
[0609] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[0610] Example 1
[0611] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0612] In recent years, dialogue systems using generative AI that mimic the personality and speaking style of a specific person have been attracting attention, but existing systems have not been able to provide realistic and personalized dialogue to reduce users' feelings of loneliness, limiting the experience.In addition, functions related to saving and secondary use of dialogue data are limited, presenting technical challenges for further enriching the user experience.
[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0614] In this invention, the server includes means for initializing modules and connecting to a database when the system starts up, means for verifying user authentication information against the database and starting a user session if authentication is successful, means for allowing a user to select a specific person, acquiring conversation data related to that person, and loading and configuring a corresponding AI model, and means for storing the dialogue data in the database in real time and providing it in a reusable format as needed, thereby enabling users to smoothly engage in real, personalized dialogue with specific people and facilitating the storage and secondary use of the dialogue data.
[0615] A "user" refers to an entity that uses a specific terminal to access the dialogue system and engage in dialogue.
[0616] A "terminal" is a device used by a user to interact, and includes a PC, smartphone, tablet, etc.
[0617] A "server" is a computer system that plays a central role in the system, performing user authentication, managing interaction data, and operating the generation AI.
[0618] "Generative AI" refers to artificial intelligence models that mimic the personality and speaking style of a specific person to provide realistic dialogue.
[0619] "Personality" refers to a person's unique character traits and distinctive style of expression.
[0620] "Speech style" refers to a particular person's tone of voice and the way they use language.
[0621] A "prompt sentence" refers to a text message entered into the generation AI, based on which the AI generates a response.
[0622] A "user session" refers to session information for managing a series of interactions while a user is logged into a system.
[0623] A "database" is a system for storing and managing various data such as dialogue data and user information.
[0624] A "module" refers to a program component for realizing a specific function of a system.
[0625] "Initialization" refers to the process by which a system or module prepares to start operation.
[0626] "Logging out" refers to a user exiting the system and ending the user session.
[0627] "Interaction Data" refers to records of messages exchanged between the user and the generating AI.
[0628] "Secondary use" refers to converting the saved dialogue data into another format, such as an audio file, or using it for other purposes.
[0629] This section describes a specific embodiment of the present invention. This system allows users to select a specific person and provides realistic dialogue through a generative AI that mimics that person's personality and speaking style. Each of the following steps details the configuration and operation of the respective hardware and software.
[0630] System Configuration
[0631] The system mainly consists of the following components:
[0632] 1. User Device
[0633] A terminal is a device with which a user interacts, and includes a PC, smartphone, tablet, etc. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0634] 2. Server
[0635] The server plays a central role in the system, performing user authentication, dialogue data management, and operation management of the generative AI. The server is composed of multiple modules (e.g., a user authentication module, a data management module, an AI engine, etc.).
[0636] System Operation
[0637] The main system operations are explained below.
[0638] 1. System startup and user authentication
[0639] When the server starts the system, all necessary modules are initialized and connected to the database.
[0640] When a user enters login information on a terminal, the terminal sends that information to the server, which checks the received authentication information against a database and starts a user session if authentication is successful. If authentication fails, it returns an error message to the terminal.
[0641] 2. Character selection and AI settings
[0642] The user selects a specific person with whom they wish to interact on the platform, and the device sends the selected person's information to the server.
[0643] Based on the received person ID, the server retrieves the conversation data related to that person from the database, loads the corresponding AI model, and sets the selected person's personality and speaking style to the AI model, completing the initialization.
[0644] 3. Start the dialogue
[0645] The user inputs a text message to start a conversation and presses the send button. The terminal sends the input message to the server.
[0646] The server passes the received message to the AI model, which generates the optimal response to that message, which is then sent back to the device, where it is displayed to the user.
[0647] 4. Response display and conversation continuation
[0648] When the user enters and sends a new message to continue the conversation, the terminal again sends the message to the server.
[0649] The server again passes the message to the AI model and generates a response, and this process repeats until the interaction is over.
[0650] 5. Storage and secondary use of conversation data
[0651] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[0652] Users can reuse their dialogue data as audio files, etc., as needed, and the server converts the data accordingly. For example, to convert text data into audio data, a Text-to-Speech (TTS) engine is used.
[0653] 6. Exiting and Logging Out
[0654] When the user wants to end the conversation, he clicks the end button, and the terminal sends a request to end the conversation to the server.
[0655] The server ends the session, stores any necessary logs, and finally logs the user out and exits the system.
[0656] Specific examples
[0657] For example, consider the process of a user logging in and selecting a particular famous person (e.g., a famous scientist) to begin a dialogue. When the user types and submits a question such as, "What do you think about the latest scientific research?", the server passes the message to an AI model, which generates an appropriate response (e.g., "The latest research is very interesting, especially its original approach.") and sends it back to the device. The device then displays this response to the user and continues the dialogue.
[0658] Prompt Sentence Examples
[0659] For example, you might input the following prompts into a generative AI model:
[0660] "What do you think about the latest scientific research?"
[0661] The prompt allows the user to ask a question to the AI model as a specific person, with the expectation that it will generate a response that mimics the way that person speaks and thinks.
[0662] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0663] Step 1: System startup and module initialization
[0664] When the server starts the system, it initializes all necessary modules (user authentication module, data management module, AI engine, etc.) and connects to the database.
[0665] Input: System startup trigger.
[0666] Processing: Module initialization, database connection establishment.
[0667] Output: System initialization complete and database connection status.
[0668] Step 2: User authentication
[0669] The user enters login information on the device and clicks the login button.
[0670] The terminal sends the entered login information (user name, password) to the server.
[0671] The server checks the received authentication information against the user information in its database, and if authentication is successful, starts a user session.
[0672] Input: Username, Password.
[0673] Processing: Database verification, session start.
[0674] Output: Session ID, authentication success or failure result.
[0675] Step 3: Selecting a person and setting AI
[0676] The user selects the specific person with whom they wish to interact on the platform.
[0677] The device sends the selected person's information (person ID and name) to the server.
[0678] Based on the received person information, the server retrieves conversation data related to that person from the database and loads the corresponding AI model.
[0679] The server sets the AI model with the selected person's personality and speaking style, completing initialization.
[0680] Input: Person ID, Name.
[0681] Processing: Acquire conversation data, load AI model, and set personality.
[0682] Output: The initialized AI model.
[0683] Step 4: Start a conversation
[0684] The user enters a text message to initiate a conversation and presses the send button.
[0685] The terminal transmits the input message to the server.
[0686] The server passes the received message to the AI model, which generates the optimal response to that message.
[0687] The generated response is sent back from the server to the terminal, which displays the response to the user.
[0688] Input: User message.
[0689] Processing: Parse the message, feed it into the AI model, and generate a response.
[0690] Output: The generated response message.
[0691] Step 5: Keep the conversation going
[0692] The user enters and sends a new message to continue the conversation.
[0693] The terminal sends a new message to the server.
[0694] The server again passes the message to the AI model, which generates a response.
[0695] This process is repeated until the conversation is over.
[0696] Input: New user message.
[0697] Processing: Parse the message, feed it into the AI model, and generate a response.
[0698] Output: The generated response message.
[0699] Step 6: Saving and Reusing Conversation Data
[0700] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[0701] The user can reuse the dialogue data as a voice file or the like.
[0702] The server performs the data conversion for this purpose, for example by using a Text-to-Speech (TTS) engine.
[0703] Input: Interaction data.
[0704] Processing: Data storage, data conversion (e.g. text to speech).
[0705] Output: Saved interaction logs, transformed data.
[0706] Step 7: End the conversation and log out
[0707] If the user wishes to end the interaction, he clicks on the End button.
[0708] The terminal sends a request to end the conversation to the server.
[0709] The server terminates the session and saves any necessary logs.
[0710] Finally, the server logs the user out and causes him to leave the system.
[0711] Input: A conversation termination request.
[0712] Processing: Session termination, log saving, logout processing.
[0713] Output: Finished sessions, saved interaction logs, logout status.
[0714] (Application example 1)
[0715] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0716] Conventional dialogue systems allow users to have a realistic dialogue experience through dialogue with a generative AI that mimics the personality and speaking style of a specific person. However, there is a need to provide a richer dialogue experience by adding entertainment and information functions. Furthermore, there is a lack of functionality that allows users to save their dialogue history and use it for later playback or learning. Technology is needed to solve these issues.
[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0718] In this invention, the server includes means for retrieving conversation data related to a selected person from a database and loading and configuring a corresponding generative AI model, means for verifying user authentication information and starting a user session if login is successful, and means for adjusting and optimizing the generative AI model for providing entertainment and information. This allows a user to enjoy entertainment and information in real time through conversations with a specific person, and saves the conversation history for later playback or learning.
[0719] A "user" is an entity that uses the system to have a conversation with a specific person.
[0720] A "specific person" is someone the user selects and a dialogue is provided through a generative AI that mimics that person's personality and speaking style.
[0721] "Personality" refers to the characteristics, traits, and speaking style that are unique to a particular person.
[0722] "Generative AI" is artificial intelligence that mimics the personality and speaking style of a specific person and generates realistic conversations with the user.
[0723] A "text message" is the text information entered by the user, and is the data that the generation AI uses to generate a response.
[0724] A "server" is a device that plays a central role in the system, performing user authentication, managing dialogue data, and operating the generation AI.
[0725] A "response" is a reply message that the generation AI generates in response to a user's text message.
[0726] A "terminal" is a device used by a user to interact, and includes a smartphone, a personal computer, etc.
[0727] "Dialogue data" is a record of the conversation between the user and the generating AI.
[0728] "Secondary use" refers to reusing saved dialogue data in a different format or for a different purpose.
[0729] "Logging out" means ending an interactive session and exiting the system by the user.
[0730] "Conversation data" is dialogue information related to a particular person.
[0731] A "database" is a system that stores and manages information such as conversation data.
[0732] An "AI model" is an algorithm or mathematical model that serves as the basis for generative AI that mimics the personality and speaking style of a specific person.
[0733] "Entertainment" refers to interactive content and functions that provide enjoyment to users.
[0734] "Information" refers to knowledge and data that users acquire through dialogue.
[0735] "Playback" refers to using the saved dialogue history to check it later.
[0736] "Logging in" is the process by which a user enters and confirms authentication information in order to access and use a system.
[0737] "Authentication information" refers to personal information required to access a system, such as a user ID and password.
[0738] A "user session" is a series of operations performed by a user from the time the user logs in until the time the user logs out.
[0739] An "error message" is a warning or guidance message that is displayed to the user when login fails.
[0740] "Tuning" is the process of configuring and optimizing a generative AI model.
[0741] "Optimization" refers to adjusting the generative AI model to improve the user's interaction experience.
[0742] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[0743] Overall system overview
[0744] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly useful for parents raising children and lonely elderly people. In addition, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement.
[0745] System Components and Operation
[0746] 1. User Device
[0747] The terminal is a device with which the user interacts, and can be a smartphone, smart glasses, or a head-mounted display. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0748] 2. Server
[0749] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[0750] Program processing flow
[0751] 1. System startup and user authentication
[0752] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0753] 2. Character selection and AI settings
[0754] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0755] 3. Start the dialogue
[0756] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[0757] 4. Response display and conversation continuation
[0758] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[0759] 5. Storage and secondary use of conversation data
[0760] The server stores the dialogue data between the user and the generated AI in a database in real time as a log. The user can reuse the dialogue data as audio files, etc., as needed, and the server converts the data for this purpose.
[0761] 6. Exiting and Logging Out
[0762] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0763] Hardware and software used
[0764] Hardware: Smartphones, smart glasses, head-mounted displays, servers
[0765] Software: Web applications, native applications, generative AI models, databases (MySQL, PostgreSQL), cloud services (AWS, Google Cloud)
[0766] Examples of specific examples and prompts
[0767] For example, consider a scenario in which a user logs into a "virtual conversation hall" and begins a conversation with a famous author. When the user types and sends a question such as "What do you think about the latest novel?", the server passes the message to an AI model, which generates an appropriate response. The generated response (e.g., "I've tried a particular theme. I'm curious to see how it resonates with you.") is sent back to the device and displayed to the user. In this way, the user can have a realistic conversation experience.
[0768] Prompt Sentence Examples
[0769] User: What do you think about the latest novel?
[0770] Server: Sends messages to the AI model and generates responses such as:
[0771] AI Model: "I've tried a particular theme and I'm curious to see how it resonates."
[0772] In this way, the present invention provides users with a realistic interaction experience with a specific person, enabling them to enjoy entertainment and information.
[0773] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0774] Step 1:
[0775] System startup and user authentication:
[0776] The server starts the system, initializes all necessary modules, and connects to the database. The user enters login information (username, password) on the terminal. The terminal sends the entered information to the server. The server compares the received login information with the database, and if authentication is successful, starts the user session. At this time, it obtains the user's authentication information from the database after hashing it and compares it with the entered information. If authentication fails, an error message is returned to the user's terminal. As output, if successful, the user session will start, and if unsuccessful, an error message will be displayed.
[0777] Step 2:
[0778] Character Selection and AI Settings:
[0779] The user selects a specific person with whom they wish to interact on the platform. The device sends the selected person's information to the server. The server uses this information to retrieve conversation data related to the selected person from a database. The server then loads the corresponding generative AI model, sets the selected person's personality and speaking style in the AI model, and completes initialization. The configured generative AI model is prepared as the output.
[0780] Step 3:
[0781] Start the conversation:
[0782] The user enters a text message to start a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to a generative AI model, which generates an optimal response to that message. In this case, the generative AI model (e.g., OpenAI GPT-3, GPT-4) takes text data as input and runs an algorithm to generate an appropriate response. The generated response is sent back from the server to the device, and the device displays it to the user. As an output, a response message is generated that is displayed on the user's device.
[0783] Step 4:
[0784] Response display and conversation continuation:
[0785] The user types and sends a new text message to continue the dialogue. The device again sends the message to the server. The server again passes the message to the generative AI model, which generates a new response. This process repeats until the dialogue ends. Each message exchange involves data processing and the generative AI generating a response based on the new user input. As an output, the continuing dialogue message is displayed on the user's device.
[0786] Step 5:
[0787] Conversation data storage and secondary use:
[0788] The server stores the dialogue data between the user and the generated AI in a database in real time. The stored dialogue data can be converted into an audio file or other format as needed for secondary use. The server uses a data conversion algorithm to convert text data into an audio file. As output, the stored dialogue data and data in a format that can be used for secondary use are generated.
[0789] Step 6:
[0790] Exit and log out:
[0791] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server. The server ends the session and saves the necessary logs. Finally, the user is logged out and exits the system. As an output, the user is safely logged out of the system.
[0792] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0793] As an embodiment of the present invention, a system is described below that combines an emotion engine with a dialogue platform using generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue. This system is realized when a user selects a specific person and sends a message to the generative AI, which recognizes the user's emotions and generates and provides a response based on them.
[0794] Overall system overview
[0795] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through generative AI. This system is designed to help reduce the user's sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. Furthermore, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly.
[0796] System Components and Operation
[0797] 1. User Device
[0798] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[0799] 2. Server
[0800] The server plays a central role in the system, performing user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[0801] 3. Generation AI
[0802] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue. The server sets the necessary parameters and generates a response based on the user's message.
[0803] 4. Emotion Engine
[0804] The emotion engine recognizes emotions from the text messages entered by the user and adjusts the generative AI's response based on those emotions. Emotion data is also stored along with the dialogue data and can be used for analysis and feedback as needed.
[0805] Program processing flow
[0806] System startup and user authentication
[0807] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[0808] Character selection and AI settings
[0809] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[0810] Start a dialogue
[0811] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to an emotion engine to recognize the user's emotion. The recognized emotion data is passed to an AI model, which generates a tailored response based on the emotion. The generated response is sent back to the device, which then displays it to the user.
[0812] Response display and conversation continuation
[0813] When the user enters and sends a new message to continue the conversation, the device again sends the message to the server, which then passes it to the emotion engine, which recognizes the user's emotion and generates an appropriate response. This process is repeated until the conversation ends.
[0814] Storage and secondary use of conversation data
[0815] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time. Users can reuse their conversation data as audio files, etc., as needed, and the server converts the data for this purpose. The server also analyzes the emotional data and uses it as feedback to improve the quality of the conversation.
[0816] Exit and log out
[0817] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[0818] Specific examples
[0819] For example, when a user logs in and selects a particular famous author to begin a dialogue, the user can type in a question like, "What do you think about the latest novel?" and send it. The server then passes the message to the emotion engine, which recognizes the user's emotion (e.g., excitement, curiosity). Based on the emotion data, the AI model generates a response (e.g., "I've tried a particular topic. I'm curious to see how it resonates with you.") and sends it back to the device. The device then displays this response to the user, continuing the dialogue. This process allows the user to have a more personal and realistic dialogue experience.
[0820] The processing flow will be explained below.
[0821] Step 1:
[0822] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, emotion engine module, and user interface module) and connects to the database.
[0823] Step 2:
[0824] The user enters a username and password into the login page on the terminal.
[0825] Step 3:
[0826] The terminal transmits the login information entered by the user to the server.
[0827] Step 4:
[0828] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is redirected to the main screen. If authentication fails, an error message is returned.
[0829] Step 5:
[0830] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[0831] Step 6:
[0832] The terminal transmits the information of the selected person to the server.
[0833] Step 7:
[0834] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking characteristics to complete initialization.
[0835] Step 8:
[0836] The user enters a text message and presses the send button to initiate a conversation.
[0837] Step 9:
[0838] The terminal transmits the input message to the server.
[0839] Step 10:
[0840] The server passes the received message to the emotion engine, which analyzes the text and recognizes the user's emotion from the message.
[0841] Step 11:
[0842] The server passes the recognized user emotion data to the AI model, which uses it to generate an appropriate response.
[0843] Step 12:
[0844] The server generates a response message adjusted based on the emotion data and transmits it to the terminal.
[0845] Step 13:
[0846] The terminal displays the response message received from the server to the user.
[0847] Step 14:
[0848] When the user enters and sends a new message to continue the conversation, the terminal sends the message to the server again.
[0849] Step 15:
[0850] The server again passes the message to the emotion engine, which recognizes the emotion and generates a response based on the recognized emotion data. This process is repeated until the interaction is completed.
[0851] Step 16:
[0852] The server stores the conversation data and emotion data between the user and the generated AI in a database in real time.
[0853] Step 17:
[0854] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[0855] Step 18:
[0856] The terminal sends a request to end the conversation to the server.
[0857] Step 19:
[0858] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[0859] Example 2
[0860] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0861] In modern society, many users experience feelings of loneliness, a major problem especially for mothers raising children and lonely elderly people. Conventional dialogue systems struggle to fully recognize users' emotions and respond appropriately, preventing users from engaging in satisfying dialogue. Furthermore, data integrity and security are not guaranteed when dialogue data is stored or reused. A system that can address these issues is needed.
[0862] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotions using an emotion engine based on a message received by the server and generating an appropriate response, means for returning the generated response to the terminal, and means for the server to store dialogue data and emotion data between the user and the generation AI and provide them in a format that can be reused. This allows the user to experience a realistic dialogue and obtain an appropriate response based on the emotion. Furthermore, since the stored dialogue data and emotion data can be reused, it is possible to improve the quality of the dialogue and increase user satisfaction.
[0863] "User" refers to a person who uses the system to interact with a specific person.
[0864] "Terminal" refers to a device such as a PC, smartphone, or tablet that a user uses to interact.
[0865] A "server" refers to a computer that plays a central role in the system and performs user authentication, dialogue data management, generative AI operation, emotion recognition, etc.
[0866] "Generative AI" refers to an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[0867] An "emotion engine" refers to a program that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[0868] "Dialogue data" refers to data that records the content of the dialogue between the user and the generating AI.
[0869] "Secondary use" refers to converting saved dialogue data and emotion data into audio files, etc., and reusing them for analysis or other purposes.
[0870] "Initialization" refers to the process of setting up a system or program so that it can operate normally.
[0871] "Personality" refers to the characteristics and traits that a particular person possesses.
[0872] "Logging in" refers to the process by which a user enters authentication information to access a system and is authenticated by the system.
[0873] A "session" refers to a series of interactions with the system from the time a user logs in until the time they log out.
[0874] This invention is a dialogue platform system that uses generative AI to mimic the personality and speaking style of a specific person and provide realistic dialogue. When a user selects a specific person and sends a message to the generative AI, the AI recognizes the user's emotions and generates and provides a response based on those emotions. This system is particularly useful for users who feel lonely, such as mothers raising children or lonely elderly people.
[0875] System Components
[0876] 1. User Device
[0877] Terminals include devices connected to the Internet, such as PCs, smartphones, tablets, etc. Users interact with these terminals.
[0878] The terminal accesses the server through a web application or a native application.
[0879] 2. Server
[0880] The server is the core of the system, and is responsible for user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[0881] The server sets the necessary parameters and generates a response according to the user's message.
[0882] 3. Generation AI
[0883] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[0884] The generative AI creates a realistic interaction experience for users based on the characteristics of the person they select.
[0885] 4. Emotion Engine
[0886] The emotion engine is responsible for recognizing emotions from the text messages entered by the user and adjusting the generative AI's response based on those emotions.
[0887] Emotional data is stored along with the dialogue data and can be used for analysis and feedback as needed.
[0888] System Operation
[0889] In this system, the user initiates a conversation using a device, and the server provides a realistic conversation via a generative AI. The overall flow of the system is shown below.
[0890] 1. User authentication: The user enters login information on the terminal, which then sends it to the server. The server checks the authentication information against a database and starts the session if authentication is successful. If authentication fails, it returns an error message.
[0891] 2. Person selection: The user selects the desired conversation partner, and the device sends this information to the server. The server retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The AI model initialization is complete.
[0892] 3. Interaction: When the user inputs a message, the device sends it to the server. The server passes the message to the emotion engine to recognize the user's emotion. It generates a response tailored based on the emotion and sends it back to the device. The device displays the response to the user.
[0893] 4. Continuous Dialogue: This process repeats as long as the user continues to input new messages. For each message, the appropriate emotion is recognized and a response is generated.
[0894] 5. Data storage and secondary use: The server stores the conversation data and emotion data between the user and the generated AI in a database. The user can reuse the conversation data as audio files, and the server converts the data for this purpose. The server also analyzes the emotion data and uses it as feedback to improve the quality of the conversation.
[0895] 6. Termination and logout: When the user finishes the conversation, the terminal sends a request to terminate the conversation to the server, which then terminates the session, saves the necessary logs, and logs out the user.
[0896] Specific examples of operation
[0897] For example, a user logs into the system, selects a particular well-known author, and begins a dialogue by typing and sending a question such as, "What do you think of his latest novel?" The server passes this message to the emotion engine, which recognizes the user's emotion (e.g., excitement or curiosity). Based on this emotion data, the generative AI generates a response such as, "I've tried a particular theme. I'm curious to see how it resonates with you," and sends it back to the device. The device displays this response to the user, allowing the user to continue the dialogue.
[0898] Prompt Sentence Examples
[0899] "What is your opinion on the latest science and technology?"
[0900] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0901] Step 1:
[0902] System startup
[0903] The server initializes all modules in the system, securing the necessary resources and establishing connections to the database.
[0904] Input: System startup instructions
[0905] Data processing: Module initialization, resource allocation
[0906] Output: Initialization complete, connection established
[0907] Step 2:
[0908] User Authentication
[0909] The user enters login information (user name, password) on the terminal and presses the send button.
[0910] The terminal transmits the entered authentication information to the server.
[0911] The server receives the input information and checks it against a database.
[0912] Input: Authentication information (username, password)
[0913] Data processing: Matching with database
[0914] Output: Authentication success or failure message
[0915] If authentication is successful, the server starts a user session and generates a corresponding session ID, otherwise it returns an error message to the terminal.
[0916] Step 3:
[0917] Character Selection
[0918] Users select the specific person on the platform they want to interact with.
[0919] The terminal transmits the information of the selected person to the server.
[0920] Based on the received information, the server retrieves relevant conversation data from the database and loads the corresponding AI model.
[0921] Input: Selected person's information
[0922] Data processing: Acquire conversation data, load AI model
[0923] Output: AI model initialization complete
[0924] Step 4:
[0925] Start a dialogue
[0926] The user enters a message and presses the send button.
[0927] The terminal transmits the input message to the server.
[0928] The server passes the received message to the emotion engine to recognize the user's emotion.
[0929] Input: User message
[0930] Data processing: Emotion recognition
[0931] Output: Emotion data
[0932] The generation AI generates a response based on the recognized emotion data, and the server sends the generated response back to the device.
[0933] Step 5:
[0934] Response Display
[0935] The terminal displays the response received from the server to the user.
[0936] Input: Generator AI response
[0937] Data processing: Displaying responses
[0938] Output: The response displayed on the user's screen
[0939] Step 6:
[0940] Ongoing dialogue
[0941] The user enters a new message and repeats steps 4 and 5 above.
[0942] Input: New user message
[0943] Data processing: emotion recognition, response generation
[0944] Output: Shows the new response
[0945] Step 7:
[0946] Saving conversation data
[0947] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time.
[0948] Input: Dialogue data, emotion data
[0949] Data processing: Saving to database
[0950] Output: Saved interaction data
[0951] Step 8:
[0952] Secondary use
[0953] The user can reuse the saved dialogue data as a voice file or the like.
[0954] The server converts the interaction data and provides it in the specified format.
[0955] Input: Saved interaction data
[0956] Data processing: Data conversion
[0957] Output: Converted data format
[0958] Step 9:
[0959] Exit and log out
[0960] If the user wishes to end the dialogue, he or she presses the end button.
[0961] The terminal sends a request to end the conversation to the server.
[0962] The server ends the session, saves any necessary logs, and logs the user out.
[0963] Input: Conversation termination request
[0964] Data processing: End session, save log
[0965] Output: Logout completed
[0966] (Application example 2)
[0967] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0968] Current dialogue systems for factory robots have difficulty achieving natural communication with humans and are unable to provide sufficient support to workers. As a result, it is difficult to improve work efficiency and safety, and there is a risk of worker fatigue and stress increasing. In addition, the robot lacks the ability to recognize workers' emotions and respond appropriately based on them, which means that there is an issue of not being able to reduce workers' psychological burden or provide sufficient support for their work.
[0969] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotion and generating a response based on that emotion, means for the server to recognize the user's emotion and adjust the response based on the emotion data, and means for the user to receive work instructions through dialogue with the robot. This not only enables natural communication when factory workers interact with the robot, but also enables appropriate responses and support that take into account the worker's emotion and fatigue. As a result, it is expected that the psychological burden on workers will be reduced and work efficiency and safety will be improved.
[0970] "User" refers to a human being who uses the system to engage in dialogue using generative AI.
[0971] A "specific person" refers to a person selected by the user whose personality and speaking style the generated AI will imitate and set up to converse with.
[0972] "Generative AI" refers to artificial intelligence programs that are designed to mimic the personality and speaking style of a specific person.
[0973] "Server" refers to the central computer system that manages the storage of interaction data, user authentication, emotion recognition, and the operation of generative AI.
[0974] An "emotion engine" refers to a module that recognizes emotions from messages entered by users and generates responses based on those emotions.
[0975] A "terminal" is a device with which a user interacts, such as a smartphone, tablet, or computer.
[0976] "Dialogue data" refers to information including messages and emotional data exchanged between the user and the generating AI.
[0977] "Secondary use" refers to reusing saved dialogue data for other purposes.
[0978] "AI model" refers to an artificial intelligence module that includes the configuration information and algorithms used by the generative AI.
[0979] "Emotion data" refers to data that recognizes a user's emotions and includes information about them.
[0980] A "robot" refers to a mechanical device used to interact with workers in factories and other places.
[0981] "Work instructions" refer to instructions regarding specific work content that a user receives through interaction with a robot.
[0982] "Login information" refers to information used for user authentication, and primarily refers to user IDs and passwords.
[0983] This invention is a system that uses a support robot to interact with factory workers. The system is designed to allow users to experience realistic interactions through a generative AI that mimics the personality and speaking style of a specific person. The specific configuration and operation for implementing this system are described below.
[0984] Overall system configuration
[0985] The system mainly consists of the following components:
[0986] 1. User Device:
[0987] The devices used are smartphones, tablets, and computers. These devices connect to the internet and communicate with the server.
[0988] 2. Server:
[0989] The Server is the central computer system that processes user input messages and generates responses using generative AI. The Server contains the following main modules:
[0990] User Authentication Module
[0991] Generative AI Module
[0992] Emotion Engine Module
[0993] Database Management Module
[0994] 3. Generation AI:
[0995] Generative AI is an artificial intelligence program that generates realistic dialogue by imitating the personality and speaking style of a specific person, and is primarily implemented using the GPT-2 model.
[0996] 4. Emotion Engine:
[0997] The emotion engine is a module that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[0998] Operation overview
[0999] The system operation is outlined below.
[1000] 1. User authentication:
[1001] The user enters login information on the terminal and sends it to the server, which verifies the information and, if authentication is successful, starts the user session.
[1002] 2. Selecting a specific person:
[1003] The user selects a specific person with whom they wish to interact, and the device transmits the selection to the server, which loads and configures the corresponding generative AI model.
[1004] 3. Start a conversation:
[1005] To initiate a dialogue, a user inputs and sends a text message, and the server passes the message to an emotion engine, which recognizes the user's emotion, generates an appropriate response, and sends it back to the device.
[1006] 4. Continuing the conversation and storing data:
[1007] As the conversation continues, new messages from the user are processed in the same way. All conversation data is stored on the server in real time and made available in a reusable format.
[1008] Usage example
[1009] For example, imagine a factory worker logs into the system and selects a specific support representative to begin a conversation. If the user types, "Please tell me how to operate this machine," the following prompt sentence is generated:
[1010] Example prompt sentence:
[1011] The following is a conversation with a supportive robot. The robot is very helpful and supportive.
[1012] Human: Please tell me how to operate this machine.
[1013] Robot (supportive):
[1014] The server uses a generative AI model to generate a response, such as, "First, turn on the machine, then follow the instructions on the control panel."
[1015] The system allows factory workers to receive specific work instructions through a realistic interactive experience, which is expected to improve efficiency and safety.
[1016] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1017] Step 1:
[1018] User authentication: The user enters login information (user ID and password) on the terminal and sends it to the server. The server checks the information against a database, and if authentication is successful, a user session begins. If authentication is successful, the server generates a session ID and returns it to the terminal. Conversely, if authentication fails, the server returns an error message.
[1019] Input: User ID, Password
[1020] Data processing: Database matching of authentication information
[1021] Output: Session ID (if successful), Error message (if unsuccessful)
[1022] Step 2:
[1023] Selecting a specific person: The user selects a specific person with whom they wish to converse. The device sends the selected person's information to the server. The server retrieves the conversation data related to that person from the database, loads and configures the generative AI model, and then configures the model with the specific person's personality and speaking style.
[1024] Input: Selected person's information
[1025] Data processing: Acquiring a database of conversation data, loading and configuring the generative AI model
[1026] Output: The configured generative AI model
[1027] Step 3:
[1028] Dialogue start: The user enters a text message to start the dialogue and sends it from the device to the server. The server passes the received message to the emotion engine, which recognizes the user's emotion. The recognized emotion data is passed to a generative AI model, which generates an appropriate response based on the emotion. The generated response is then sent back to the device and displayed to the user.
[1029] Input: User's text message
[1030] Data processing: Emotion recognition and response generation using generative AI
[1031] Output: Response message
[1032] Step 4:
[1033] Continuing the dialogue and saving data: Every time the user inputs a new message to continue the dialogue, the device sends the message to the server. Emotion recognition is performed in the same way, and a response is generated based on the emotional data. This series of processes is then repeated. Furthermore, the server saves all dialogue data in a database.
[1034] Input: The user's new text message
[1035] Data processing: Emotion recognition, response generation using generative AI, and dialogue data storage
[1036] Output: Response message, saved interaction data
[1037] Step 5:
[1038] Ending the conversation and logging out: When the user wants to end the conversation, he clicks the end button and the terminal sends a request to end the conversation to the server. The server ends the conversation session and saves the log information. After that, the user is logged out and exits the system.
[1039] Input: Conversation termination request
[1040] Data processing: End session, save log information
[1041] Output: Logged out of the system
[1042] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1043] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1044] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1045] [Third embodiment]
[1046] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1047] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1048] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1049] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1050] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1051] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1052] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1053] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1054] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1055] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1056] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1057] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1058] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[1059] Overall system overview
[1060] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. In addition, the conversation data from the system is saved and can be reused, reducing the risk of copyright infringement.
[1061] System Components and Operation
[1062] 1. User Device
[1063] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1064] 2. Server
[1065] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[1066] Program processing flow
[1067] System startup and user authentication
[1068] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1069] Character selection and AI settings
[1070] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1071] Start a dialogue
[1072] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[1073] Response display and conversation continuation
[1074] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1075] Storage and secondary use of conversation data
[1076] The server stores the conversation data between the user and the generated AI in a database in real time as a log. Users can reuse their conversation data as needed, such as creating audio files, and the server converts the data for this purpose.
[1077] Exit and log out
[1078] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1079] Specific examples
[1080] For example, consider the process of a user logging in, selecting a particular famous author, and starting a dialogue. When the user types and sends a question such as, "What do you think of the latest novel?", the server passes the message to an AI model, which generates an appropriate response (e.g., "I've tried a particular subject. I'm curious to know how it resonates with you.") and sends it back to the device. The device displays this response to the user, and the dialogue continues. This series of processes allows the user to have a realistic dialogue experience.
[1081] The processing flow will be explained below.
[1082] Step 1:
[1083] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, user interface module) and connects to the database.
[1084] Step 2:
[1085] The user enters a username and password into the login page on the terminal.
[1086] Step 3:
[1087] The terminal transmits the login information entered by the user to the server.
[1088] Step 4:
[1089] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is taken to the main screen.
[1090] Step 5:
[1091] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[1092] Step 6:
[1093] The terminal transmits the information of the selected person to the server.
[1094] Step 7:
[1095] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking style to complete the initialization.
[1096] Step 8:
[1097] The user enters a text message and presses the send button to initiate a conversation.
[1098] Step 9:
[1099] The terminal transmits the input message to the server.
[1100] Step 10:
[1101] The server passes the received message to the AI model, which generates the optimal response to that message.
[1102] Step 11:
[1103] The server sends the generated response message to the terminal.
[1104] Step 12:
[1105] The terminal displays the response message received from the server to the user.
[1106] Step 13:
[1107] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1108] Step 14:
[1109] The server stores the interaction data between the user and the generated AI in a database in real time.
[1110] Step 15:
[1111] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[1112] Step 16:
[1113] The terminal sends a request to end the conversation to the server.
[1114] Step 17:
[1115] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[1116] Example 1
[1117] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1118] In recent years, dialogue systems using generative AI that mimic the personality and speaking style of a specific person have been attracting attention, but existing systems have not been able to provide realistic and personalized dialogue to reduce users' feelings of loneliness, limiting the experience.In addition, functions related to saving and secondary use of dialogue data are limited, presenting technical challenges for further enriching the user experience.
[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1120] In this invention, the server includes means for initializing modules and connecting to a database when the system starts up, means for verifying user authentication information against the database and starting a user session if authentication is successful, means for allowing a user to select a specific person, acquiring conversation data related to that person, and loading and configuring a corresponding AI model, and means for storing the dialogue data in the database in real time and providing it in a reusable format as needed, thereby enabling users to smoothly engage in real, personalized dialogue with specific people and facilitating the storage and secondary use of the dialogue data.
[1121] A "user" refers to an entity that uses a specific terminal to access the dialogue system and engage in dialogue.
[1122] A "terminal" is a device used by a user to interact, and includes a PC, smartphone, tablet, etc.
[1123] A "server" is a computer system that plays a central role in the system, performing user authentication, managing interaction data, and operating the generation AI.
[1124] "Generative AI" refers to artificial intelligence models that mimic the personality and speaking style of a specific person to provide realistic dialogue.
[1125] "Personality" refers to a person's unique character traits and distinctive style of expression.
[1126] "Speech style" refers to a particular person's tone of voice and the way they use language.
[1127] A "prompt sentence" refers to a text message entered into the generation AI, based on which the AI generates a response.
[1128] A "user session" refers to session information for managing a series of interactions while a user is logged into a system.
[1129] A "database" is a system for storing and managing various data such as dialogue data and user information.
[1130] A "module" refers to a program component for realizing a specific function of a system.
[1131] "Initialization" refers to the process by which a system or module prepares to start operation.
[1132] "Logging out" refers to a user exiting the system and ending the user session.
[1133] "Interaction Data" refers to records of messages exchanged between the user and the generating AI.
[1134] "Secondary use" refers to converting the saved dialogue data into another format, such as an audio file, or using it for other purposes.
[1135] This section describes a specific embodiment of the present invention. This system allows users to select a specific person and provides realistic dialogue through a generative AI that mimics that person's personality and speaking style. Each of the following steps details the configuration and operation of the respective hardware and software.
[1136] System Configuration
[1137] The system mainly consists of the following components:
[1138] 1. User Device
[1139] A terminal is a device with which a user interacts, and includes a PC, smartphone, tablet, etc. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1140] 2. Server
[1141] The server plays a central role in the system, performing user authentication, dialogue data management, and operation management of the generative AI. The server is composed of multiple modules (e.g., a user authentication module, a data management module, an AI engine, etc.).
[1142] System Operation
[1143] The main system operations are explained below.
[1144] 1. System startup and user authentication
[1145] When the server starts the system, all necessary modules are initialized and connected to the database.
[1146] When a user enters login information on a terminal, the terminal sends that information to the server, which checks the received authentication information against a database and starts a user session if authentication is successful. If authentication fails, it returns an error message to the terminal.
[1147] 2. Character selection and AI settings
[1148] The user selects a specific person with whom they wish to interact on the platform, and the device sends the selected person's information to the server.
[1149] Based on the received person ID, the server retrieves the conversation data related to that person from the database, loads the corresponding AI model, and sets the selected person's personality and speaking style to the AI model, completing the initialization.
[1150] 3. Start the dialogue
[1151] The user inputs a text message to start a conversation and presses the send button. The terminal sends the input message to the server.
[1152] The server passes the received message to the AI model, which generates the optimal response to that message, which is then sent back to the device, where it is displayed to the user.
[1153] 4. Response display and conversation continuation
[1154] When the user enters and sends a new message to continue the conversation, the terminal again sends the message to the server.
[1155] The server again passes the message to the AI model and generates a response, and this process repeats until the interaction is over.
[1156] 5. Storage and secondary use of conversation data
[1157] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[1158] Users can reuse their dialogue data as audio files, etc., as needed, and the server converts the data accordingly. For example, to convert text data into audio data, a Text-to-Speech (TTS) engine is used.
[1159] 6. Exiting and Logging Out
[1160] When the user wants to end the conversation, he clicks the end button, and the terminal sends a request to end the conversation to the server.
[1161] The server ends the session, stores any necessary logs, and finally logs the user out and exits the system.
[1162] Specific examples
[1163] For example, consider the process of a user logging in and selecting a particular famous person (e.g., a famous scientist) to begin a dialogue. When the user types and submits a question such as, "What do you think about the latest scientific research?", the server passes the message to an AI model, which generates an appropriate response (e.g., "The latest research is very interesting, especially its original approach.") and sends it back to the device. The device then displays this response to the user and continues the dialogue.
[1164] Prompt Sentence Examples
[1165] For example, you might input the following prompts into a generative AI model:
[1166] "What do you think about the latest scientific research?"
[1167] The prompt allows the user to ask a question to the AI model as a specific person, with the expectation that it will generate a response that mimics the way that person speaks and thinks.
[1168] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1169] Step 1: System startup and module initialization
[1170] When the server starts the system, it initializes all necessary modules (user authentication module, data management module, AI engine, etc.) and connects to the database.
[1171] Input: System startup trigger.
[1172] Processing: Module initialization, database connection establishment.
[1173] Output: System initialization complete and database connection status.
[1174] Step 2: User authentication
[1175] The user enters login information on the device and clicks the login button.
[1176] The terminal sends the entered login information (user name, password) to the server.
[1177] The server checks the received authentication information against the user information in its database, and if authentication is successful, starts a user session.
[1178] Input: Username, Password.
[1179] Processing: Database verification, session start.
[1180] Output: Session ID, authentication success or failure result.
[1181] Step 3: Selecting a person and setting AI
[1182] The user selects the specific person with whom they wish to interact on the platform.
[1183] The device sends the selected person's information (person ID and name) to the server.
[1184] Based on the received person information, the server retrieves conversation data related to that person from the database and loads the corresponding AI model.
[1185] The server sets the AI model with the selected person's personality and speaking style, completing initialization.
[1186] Input: Person ID, Name.
[1187] Processing: Acquire conversation data, load AI model, and set personality.
[1188] Output: The initialized AI model.
[1189] Step 4: Start a conversation
[1190] The user enters a text message to initiate a conversation and presses the send button.
[1191] The terminal transmits the input message to the server.
[1192] The server passes the received message to the AI model, which generates the optimal response to that message.
[1193] The generated response is sent back from the server to the terminal, which displays the response to the user.
[1194] Input: User message.
[1195] Processing: Parse the message, feed it into the AI model, and generate a response.
[1196] Output: The generated response message.
[1197] Step 5: Keep the conversation going
[1198] The user enters and sends a new message to continue the conversation.
[1199] The terminal sends a new message to the server.
[1200] The server again passes the message to the AI model, which generates a response.
[1201] This process is repeated until the conversation is over.
[1202] Input: New user message.
[1203] Processing: Parse the message, feed it into the AI model, and generate a response.
[1204] Output: The generated response message.
[1205] Step 6: Saving and Reusing Conversation Data
[1206] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[1207] The user can reuse the dialogue data as a voice file or the like.
[1208] The server performs the data conversion for this purpose, for example by using a Text-to-Speech (TTS) engine.
[1209] Input: Interaction data.
[1210] Processing: Data storage, data conversion (e.g. text to speech).
[1211] Output: Saved interaction logs, transformed data.
[1212] Step 7: End the conversation and log out
[1213] If the user wishes to end the interaction, he clicks on the End button.
[1214] The terminal sends a request to end the conversation to the server.
[1215] The server terminates the session and saves any necessary logs.
[1216] Finally, the server logs the user out and causes him to leave the system.
[1217] Input: A conversation termination request.
[1218] Processing: Session termination, log saving, logout processing.
[1219] Output: Finished sessions, saved interaction logs, logout status.
[1220] (Application example 1)
[1221] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1222] Conventional dialogue systems allow users to have a realistic dialogue experience through dialogue with a generative AI that mimics the personality and speaking style of a specific person. However, there is a need to provide a richer dialogue experience by adding entertainment and information functions. Furthermore, there is a lack of functionality that allows users to save their dialogue history and use it for later playback or learning. Technology is needed to solve these issues.
[1223] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1224] In this invention, the server includes means for retrieving conversation data related to a selected person from a database and loading and configuring a corresponding generative AI model, means for verifying user authentication information and starting a user session if login is successful, and means for adjusting and optimizing the generative AI model for providing entertainment and information. This allows a user to enjoy entertainment and information in real time through conversations with a specific person, and saves the conversation history for later playback or learning.
[1225] A "user" is an entity that uses the system to have a conversation with a specific person.
[1226] A "specific person" is someone the user selects and a dialogue is provided through a generative AI that mimics that person's personality and speaking style.
[1227] "Personality" refers to the characteristics, traits, and speaking style that are unique to a particular person.
[1228] "Generative AI" is artificial intelligence that mimics the personality and speaking style of a specific person and generates realistic conversations with the user.
[1229] A "text message" is the text information entered by the user, and is the data that the generation AI uses to generate a response.
[1230] A "server" is a device that plays a central role in the system, performing user authentication, managing dialogue data, and operating the generation AI.
[1231] A "response" is a reply message that the generation AI generates in response to a user's text message.
[1232] A "terminal" is a device used by a user to interact, and includes a smartphone, a personal computer, etc.
[1233] "Dialogue data" is a record of the conversation between the user and the generating AI.
[1234] "Secondary use" refers to reusing saved dialogue data in a different format or for a different purpose.
[1235] "Logging out" means ending an interactive session and exiting the system by the user.
[1236] "Conversation data" is dialogue information related to a particular person.
[1237] A "database" is a system that stores and manages information such as conversation data.
[1238] An "AI model" is an algorithm or mathematical model that serves as the basis for generative AI that mimics the personality and speaking style of a specific person.
[1239] "Entertainment" refers to interactive content and functions that provide enjoyment to users.
[1240] "Information" refers to knowledge and data that users acquire through dialogue.
[1241] "Playback" refers to using the saved dialogue history to check it later.
[1242] "Logging in" is the process by which a user enters and confirms authentication information in order to access and use a system.
[1243] "Authentication information" refers to personal information required to access a system, such as a user ID and password.
[1244] A "user session" is a series of operations performed by a user from the time the user logs in until the time the user logs out.
[1245] An "error message" is a warning or guidance message that is displayed to the user when login fails.
[1246] "Tuning" is the process of configuring and optimizing a generative AI model.
[1247] "Optimization" refers to adjusting the generative AI model to improve the user's interaction experience.
[1248] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[1249] Overall system overview
[1250] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly useful for parents raising children and lonely elderly people. In addition, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement.
[1251] System Components and Operation
[1252] 1. User Device
[1253] The terminal is a device with which the user interacts, and can be a smartphone, smart glasses, or a head-mounted display. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1254] 2. Server
[1255] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[1256] Program processing flow
[1257] 1. System startup and user authentication
[1258] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1259] 2. Character selection and AI settings
[1260] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1261] 3. Start the dialogue
[1262] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[1263] 4. Response display and conversation continuation
[1264] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1265] 5. Storage and secondary use of conversation data
[1266] The server stores the dialogue data between the user and the generated AI in a database in real time as a log. The user can reuse the dialogue data as audio files, etc., as needed, and the server converts the data for this purpose.
[1267] 6. Exiting and Logging Out
[1268] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1269] Hardware and software used
[1270] Hardware: Smartphones, smart glasses, head-mounted displays, servers
[1271] Software: Web applications, native applications, generative AI models, databases (MySQL, PostgreSQL), cloud services (AWS, Google Cloud)
[1272] Examples of specific examples and prompts
[1273] For example, consider a scenario in which a user logs into a "virtual conversation hall" and begins a conversation with a famous author. When the user types and sends a question such as "What do you think about the latest novel?", the server passes the message to an AI model, which generates an appropriate response. The generated response (e.g., "I've tried a particular theme. I'm curious to see how it resonates with you.") is sent back to the device and displayed to the user. In this way, the user can have a realistic conversation experience.
[1274] Prompt Sentence Examples
[1275] User: What do you think about the latest novel?
[1276] Server: Sends messages to the AI model and generates responses such as:
[1277] AI Model: "I've tried a particular theme and I'm curious to see how it resonates."
[1278] In this way, the present invention provides users with a realistic interaction experience with a specific person, enabling them to enjoy entertainment and information.
[1279] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1280] Step 1:
[1281] System startup and user authentication:
[1282] The server starts the system, initializes all necessary modules, and connects to the database. The user enters login information (username, password) on the terminal. The terminal sends the entered information to the server. The server compares the received login information with the database, and if authentication is successful, starts the user session. At this time, it obtains the user's authentication information from the database after hashing it and compares it with the entered information. If authentication fails, an error message is returned to the user's terminal. As output, if successful, the user session will start, and if unsuccessful, an error message will be displayed.
[1283] Step 2:
[1284] Character Selection and AI Settings:
[1285] The user selects a specific person with whom they wish to interact on the platform. The device sends the selected person's information to the server. The server uses this information to retrieve conversation data related to the selected person from a database. The server then loads the corresponding generative AI model, sets the selected person's personality and speaking style in the AI model, and completes initialization. The configured generative AI model is prepared as the output.
[1286] Step 3:
[1287] Start the conversation:
[1288] The user enters a text message to start a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to a generative AI model, which generates an optimal response to that message. In this case, the generative AI model (e.g., OpenAI GPT-3, GPT-4) takes text data as input and runs an algorithm to generate an appropriate response. The generated response is sent back from the server to the device, and the device displays it to the user. As an output, a response message is generated that is displayed on the user's device.
[1289] Step 4:
[1290] Response display and conversation continuation:
[1291] The user types and sends a new text message to continue the dialogue. The device again sends the message to the server. The server again passes the message to the generative AI model, which generates a new response. This process repeats until the dialogue ends. Each message exchange involves data processing and the generative AI generating a response based on the new user input. As an output, the continuing dialogue message is displayed on the user's device.
[1292] Step 5:
[1293] Conversation data storage and secondary use:
[1294] The server stores the dialogue data between the user and the generated AI in a database in real time. The stored dialogue data can be converted into an audio file or other format as needed for secondary use. The server uses a data conversion algorithm to convert text data into an audio file. As output, the stored dialogue data and data in a format that can be used for secondary use are generated.
[1295] Step 6:
[1296] Exit and log out:
[1297] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server. The server ends the session and saves the necessary logs. Finally, the user is logged out and exits the system. As an output, the user is safely logged out of the system.
[1298] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1299] As an embodiment of the present invention, a system is described below that combines an emotion engine with a dialogue platform using generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue. This system is realized when a user selects a specific person and sends a message to the generative AI, which recognizes the user's emotions and generates and provides a response based on them.
[1300] Overall system overview
[1301] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through generative AI. This system is designed to help reduce the user's sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. Furthermore, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly.
[1302] System Components and Operation
[1303] 1. User Device
[1304] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1305] 2. Server
[1306] The server plays a central role in the system, performing user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[1307] 3. Generation AI
[1308] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue. The server sets the necessary parameters and generates a response based on the user's message.
[1309] 4. Emotion Engine
[1310] The emotion engine recognizes emotions from the text messages entered by the user and adjusts the generative AI's response based on those emotions. Emotion data is also stored along with the dialogue data and can be used for analysis and feedback as needed.
[1311] Program processing flow
[1312] System startup and user authentication
[1313] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1314] Character selection and AI settings
[1315] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1316] Start a dialogue
[1317] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to an emotion engine to recognize the user's emotion. The recognized emotion data is passed to an AI model, which generates a tailored response based on the emotion. The generated response is sent back to the device, which then displays it to the user.
[1318] Response display and conversation continuation
[1319] When the user enters and sends a new message to continue the conversation, the device again sends the message to the server, which then passes it to the emotion engine, which recognizes the user's emotion and generates an appropriate response. This process is repeated until the conversation ends.
[1320] Storage and secondary use of conversation data
[1321] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time. Users can reuse their conversation data as audio files, etc., as needed, and the server converts the data for this purpose. The server also analyzes the emotional data and uses it as feedback to improve the quality of the conversation.
[1322] Exit and log out
[1323] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1324] Specific examples
[1325] For example, when a user logs in and selects a particular famous author to begin a dialogue, the user can type in a question like, "What do you think about the latest novel?" and send it. The server then passes the message to the emotion engine, which recognizes the user's emotion (e.g., excitement, curiosity). Based on the emotion data, the AI model generates a response (e.g., "I've tried a particular topic. I'm curious to see how it resonates with you.") and sends it back to the device. The device then displays this response to the user, continuing the dialogue. This process allows the user to have a more personal and realistic dialogue experience.
[1326] The processing flow will be explained below.
[1327] Step 1:
[1328] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, emotion engine module, and user interface module) and connects to the database.
[1329] Step 2:
[1330] The user enters a username and password into the login page on the terminal.
[1331] Step 3:
[1332] The terminal transmits the login information entered by the user to the server.
[1333] Step 4:
[1334] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is redirected to the main screen. If authentication fails, an error message is returned.
[1335] Step 5:
[1336] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[1337] Step 6:
[1338] The terminal transmits the information of the selected person to the server.
[1339] Step 7:
[1340] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking characteristics to complete initialization.
[1341] Step 8:
[1342] The user enters a text message and presses the send button to initiate a conversation.
[1343] Step 9:
[1344] The terminal transmits the input message to the server.
[1345] Step 10:
[1346] The server passes the received message to the emotion engine, which analyzes the text and recognizes the user's emotion from the message.
[1347] Step 11:
[1348] The server passes the recognized user emotion data to the AI model, which uses it to generate an appropriate response.
[1349] Step 12:
[1350] The server generates a response message adjusted based on the emotion data and transmits it to the terminal.
[1351] Step 13:
[1352] The terminal displays the response message received from the server to the user.
[1353] Step 14:
[1354] When the user enters and sends a new message to continue the conversation, the terminal sends the message to the server again.
[1355] Step 15:
[1356] The server again passes the message to the emotion engine, which recognizes the emotion and generates a response based on the recognized emotion data. This process is repeated until the interaction is completed.
[1357] Step 16:
[1358] The server stores the conversation data and emotion data between the user and the generated AI in a database in real time.
[1359] Step 17:
[1360] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[1361] Step 18:
[1362] The terminal sends a request to end the conversation to the server.
[1363] Step 19:
[1364] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[1365] Example 2
[1366] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1367] In modern society, many users experience feelings of loneliness, a major problem especially for mothers raising children and lonely elderly people. Conventional dialogue systems struggle to fully recognize users' emotions and respond appropriately, preventing users from engaging in satisfying dialogue. Furthermore, data integrity and security are not guaranteed when dialogue data is stored or reused. A system that can address these issues is needed.
[1368] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotions using an emotion engine based on a message received by the server and generating an appropriate response, means for returning the generated response to the terminal, and means for the server to store dialogue data and emotion data between the user and the generation AI and provide them in a format that can be reused. This allows the user to experience a realistic dialogue and obtain an appropriate response based on the emotion. Furthermore, since the stored dialogue data and emotion data can be reused, it is possible to improve the quality of the dialogue and increase user satisfaction.
[1369] "User" refers to a person who uses the system to interact with a specific person.
[1370] "Terminal" refers to a device such as a PC, smartphone, or tablet that a user uses to interact.
[1371] A "server" refers to a computer that plays a central role in the system and performs user authentication, dialogue data management, generative AI operation, emotion recognition, etc.
[1372] "Generative AI" refers to an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[1373] An "emotion engine" refers to a program that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[1374] "Dialogue data" refers to data that records the content of the dialogue between the user and the generating AI.
[1375] "Secondary use" refers to converting saved dialogue data and emotion data into audio files, etc., and reusing them for analysis or other purposes.
[1376] "Initialization" refers to the process of setting up a system or program so that it can operate normally.
[1377] "Personality" refers to the characteristics and traits that a particular person possesses.
[1378] "Logging in" refers to the process by which a user enters authentication information to access a system and is authenticated by the system.
[1379] A "session" refers to a series of interactions with the system from the time a user logs in until the time they log out.
[1380] This invention is a dialogue platform system that uses generative AI to mimic the personality and speaking style of a specific person and provide realistic dialogue. When a user selects a specific person and sends a message to the generative AI, the AI recognizes the user's emotions and generates and provides a response based on those emotions. This system is particularly useful for users who feel lonely, such as mothers raising children or lonely elderly people.
[1381] System Components
[1382] 1. User Device
[1383] Terminals include devices connected to the Internet, such as PCs, smartphones, tablets, etc. Users interact with these terminals.
[1384] The terminal accesses the server through a web application or a native application.
[1385] 2. Server
[1386] The server is the core of the system, and is responsible for user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[1387] The server sets the necessary parameters and generates a response according to the user's message.
[1388] 3. Generation AI
[1389] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[1390] The generative AI creates a realistic interaction experience for users based on the characteristics of the person they select.
[1391] 4. Emotion Engine
[1392] The emotion engine is responsible for recognizing emotions from the text messages entered by the user and adjusting the generative AI's response based on those emotions.
[1393] Emotional data is stored along with the dialogue data and can be used for analysis and feedback as needed.
[1394] System Operation
[1395] In this system, the user initiates a conversation using a device, and the server provides a realistic conversation via a generative AI. The overall flow of the system is shown below.
[1396] 1. User authentication: The user enters login information on the terminal, which then sends it to the server. The server checks the authentication information against a database and starts the session if authentication is successful. If authentication fails, it returns an error message.
[1397] 2. Person selection: The user selects the desired conversation partner, and the device sends this information to the server. The server retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The AI model initialization is complete.
[1398] 3. Interaction: When the user inputs a message, the device sends it to the server. The server passes the message to the emotion engine to recognize the user's emotion. It generates a response tailored based on the emotion and sends it back to the device. The device displays the response to the user.
[1399] 4. Continuous Dialogue: This process repeats as long as the user continues to input new messages. For each message, the appropriate emotion is recognized and a response is generated.
[1400] 5. Data storage and secondary use: The server stores the conversation data and emotion data between the user and the generated AI in a database. The user can reuse the conversation data as audio files, and the server converts the data for this purpose. The server also analyzes the emotion data and uses it as feedback to improve the quality of the conversation.
[1401] 6. Termination and logout: When the user finishes the conversation, the terminal sends a request to terminate the conversation to the server, which then terminates the session, saves the necessary logs, and logs out the user.
[1402] Specific examples of operation
[1403] For example, a user logs into the system, selects a particular well-known author, and begins a dialogue by typing and sending a question such as, "What do you think of his latest novel?" The server passes this message to the emotion engine, which recognizes the user's emotion (e.g., excitement or curiosity). Based on this emotion data, the generative AI generates a response such as, "I've tried a particular theme. I'm curious to see how it resonates with you," and sends it back to the device. The device displays this response to the user, allowing the user to continue the dialogue.
[1404] Prompt Sentence Examples
[1405] "What is your opinion on the latest science and technology?"
[1406] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1407] Step 1:
[1408] System startup
[1409] The server initializes all modules in the system, securing the necessary resources and establishing connections to the database.
[1410] Input: System startup instructions
[1411] Data processing: Module initialization, resource allocation
[1412] Output: Initialization complete, connection established
[1413] Step 2:
[1414] User Authentication
[1415] The user enters login information (user name, password) on the terminal and presses the send button.
[1416] The terminal transmits the entered authentication information to the server.
[1417] The server receives the input information and checks it against a database.
[1418] Input: Authentication information (username, password)
[1419] Data processing: Matching with database
[1420] Output: Authentication success or failure message
[1421] If authentication is successful, the server starts a user session and generates a corresponding session ID, otherwise it returns an error message to the terminal.
[1422] Step 3:
[1423] Character Selection
[1424] Users select the specific person on the platform they want to interact with.
[1425] The terminal transmits the information of the selected person to the server.
[1426] Based on the received information, the server retrieves relevant conversation data from the database and loads the corresponding AI model.
[1427] Input: Selected person's information
[1428] Data processing: Acquire conversation data, load AI model
[1429] Output: AI model initialization complete
[1430] Step 4:
[1431] Start a dialogue
[1432] The user enters a message and presses the send button.
[1433] The terminal transmits the input message to the server.
[1434] The server passes the received message to the emotion engine to recognize the user's emotion.
[1435] Input: User message
[1436] Data processing: Emotion recognition
[1437] Output: Emotion data
[1438] The generation AI generates a response based on the recognized emotion data, and the server sends the generated response back to the device.
[1439] Step 5:
[1440] Response Display
[1441] The terminal displays the response received from the server to the user.
[1442] Input: Generator AI response
[1443] Data processing: Displaying responses
[1444] Output: The response displayed on the user's screen
[1445] Step 6:
[1446] Ongoing dialogue
[1447] The user enters a new message and repeats steps 4 and 5 above.
[1448] Input: New user message
[1449] Data processing: emotion recognition, response generation
[1450] Output: Shows the new response
[1451] Step 7:
[1452] Saving conversation data
[1453] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time.
[1454] Input: Dialogue data, emotion data
[1455] Data processing: Saving to database
[1456] Output: Saved interaction data
[1457] Step 8:
[1458] Secondary use
[1459] The user can reuse the saved dialogue data as a voice file or the like.
[1460] The server converts the interaction data and provides it in the specified format.
[1461] Input: Saved interaction data
[1462] Data processing: Data conversion
[1463] Output: Converted data format
[1464] Step 9:
[1465] Exit and log out
[1466] If the user wishes to end the dialogue, he or she presses the end button.
[1467] The terminal sends a request to end the conversation to the server.
[1468] The server ends the session, saves any necessary logs, and logs the user out.
[1469] Input: Conversation termination request
[1470] Data processing: End session, save log
[1471] Output: Logout completed
[1472] (Application example 2)
[1473] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1474] Current dialogue systems for factory robots have difficulty achieving natural communication with humans and are unable to provide sufficient support to workers. As a result, it is difficult to improve work efficiency and safety, and there is a risk of worker fatigue and stress increasing. In addition, the robot lacks the ability to recognize workers' emotions and respond appropriately based on them, which means that there is an issue of not being able to reduce workers' psychological burden or provide sufficient support for their work.
[1475] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotion and generating a response based on that emotion, means for the server to recognize the user's emotion and adjust the response based on the emotion data, and means for the user to receive work instructions through dialogue with the robot. This not only enables natural communication when factory workers interact with the robot, but also enables appropriate responses and support that take into account the worker's emotion and fatigue. As a result, it is expected that the psychological burden on workers will be reduced and work efficiency and safety will be improved.
[1476] "User" refers to a human being who uses the system to engage in dialogue using generative AI.
[1477] A "specific person" refers to a person selected by the user whose personality and speaking style the generated AI will imitate and set up to converse with.
[1478] "Generative AI" refers to artificial intelligence programs that are designed to mimic the personality and speaking style of a specific person.
[1479] "Server" refers to the central computer system that manages the storage of interaction data, user authentication, emotion recognition, and the operation of generative AI.
[1480] An "emotion engine" refers to a module that recognizes emotions from messages entered by users and generates responses based on those emotions.
[1481] A "terminal" is a device with which a user interacts, such as a smartphone, tablet, or computer.
[1482] "Dialogue data" refers to information including messages and emotional data exchanged between the user and the generating AI.
[1483] "Secondary use" refers to reusing saved dialogue data for other purposes.
[1484] "AI model" refers to an artificial intelligence module that includes the configuration information and algorithms used by the generative AI.
[1485] "Emotion data" refers to data that recognizes a user's emotions and includes information about them.
[1486] A "robot" refers to a mechanical device used to interact with workers in factories and other places.
[1487] "Work instructions" refer to instructions regarding specific work content that a user receives through interaction with a robot.
[1488] "Login information" refers to information used for user authentication, and primarily refers to user IDs and passwords.
[1489] This invention is a system that uses a support robot to interact with factory workers. The system is designed to allow users to experience realistic interactions through a generative AI that mimics the personality and speaking style of a specific person. The specific configuration and operation for implementing this system are described below.
[1490] Overall system configuration
[1491] The system mainly consists of the following components:
[1492] 1. User Device:
[1493] The devices used are smartphones, tablets, and computers. These devices connect to the internet and communicate with the server.
[1494] 2. Server:
[1495] The Server is the central computer system that processes user input messages and generates responses using generative AI. The Server contains the following main modules:
[1496] User Authentication Module
[1497] Generative AI Module
[1498] Emotion Engine Module
[1499] Database Management Module
[1500] 3. Generation AI:
[1501] Generative AI is an artificial intelligence program that generates realistic dialogue by imitating the personality and speaking style of a specific person, and is primarily implemented using the GPT-2 model.
[1502] 4. Emotion Engine:
[1503] The emotion engine is a module that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[1504] Operation overview
[1505] The system operation is outlined below.
[1506] 1. User authentication:
[1507] The user enters login information on the terminal and sends it to the server, which verifies the information and, if authentication is successful, starts the user session.
[1508] 2. Selecting a specific person:
[1509] The user selects a specific person with whom they wish to interact, and the device transmits the selection to the server, which loads and configures the corresponding generative AI model.
[1510] 3. Start a conversation:
[1511] To initiate a dialogue, a user inputs and sends a text message, and the server passes the message to an emotion engine, which recognizes the user's emotion, generates an appropriate response, and sends it back to the device.
[1512] 4. Continuing the conversation and storing data:
[1513] As the conversation continues, new messages from the user are processed in the same way. All conversation data is stored on the server in real time and made available in a reusable format.
[1514] Usage example
[1515] For example, imagine a factory worker logs into the system and selects a specific support representative to begin a conversation. If the user types, "Please tell me how to operate this machine," the following prompt sentence is generated:
[1516] Example prompt sentence:
[1517] The following is a conversation with a supportive robot. The robot is very helpful and supportive.
[1518] Human: Please tell me how to operate this machine.
[1519] Robot (supportive):
[1520] The server uses a generative AI model to generate a response, such as, "First, turn on the machine, then follow the instructions on the control panel."
[1521] The system allows factory workers to receive specific work instructions through a realistic interactive experience, which is expected to improve efficiency and safety.
[1522] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1523] Step 1:
[1524] User authentication: The user enters login information (user ID and password) on the terminal and sends it to the server. The server checks the information against a database, and if authentication is successful, a user session begins. If authentication is successful, the server generates a session ID and returns it to the terminal. Conversely, if authentication fails, the server returns an error message.
[1525] Input: User ID, Password
[1526] Data processing: Database matching of authentication information
[1527] Output: Session ID (if successful), Error message (if unsuccessful)
[1528] Step 2:
[1529] Selecting a specific person: The user selects a specific person with whom they wish to converse. The device sends the selected person's information to the server. The server retrieves the conversation data related to that person from the database, loads and configures the generative AI model, and then configures the model with the specific person's personality and speaking style.
[1530] Input: Selected person's information
[1531] Data processing: Acquiring a database of conversation data, loading and configuring the generative AI model
[1532] Output: The configured generative AI model
[1533] Step 3:
[1534] Dialogue start: The user enters a text message to start the dialogue and sends it from the device to the server. The server passes the received message to the emotion engine, which recognizes the user's emotion. The recognized emotion data is passed to a generative AI model, which generates an appropriate response based on the emotion. The generated response is then sent back to the device and displayed to the user.
[1535] Input: User's text message
[1536] Data processing: Emotion recognition and response generation using generative AI
[1537] Output: Response message
[1538] Step 4:
[1539] Continuing the dialogue and saving data: Every time the user inputs a new message to continue the dialogue, the device sends the message to the server. Emotion recognition is performed in the same way, and a response is generated based on the emotional data. This series of processes is then repeated. Furthermore, the server saves all dialogue data in a database.
[1540] Input: The user's new text message
[1541] Data processing: Emotion recognition, response generation using generative AI, and dialogue data storage
[1542] Output: Response message, saved interaction data
[1543] Step 5:
[1544] Ending the conversation and logging out: When the user wants to end the conversation, he clicks the end button and the terminal sends a request to end the conversation to the server. The server ends the conversation session and saves the log information. After that, the user is logged out and exits the system.
[1545] Input: Conversation termination request
[1546] Data processing: End session, save log information
[1547] Output: Logged out of the system
[1548] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1549] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1550] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1551] [Fourth embodiment]
[1552] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1553] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1554] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1555] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1556] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1558] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1559] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1560] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1561] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1562] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1563] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1564] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1565] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[1566] Overall system overview
[1567] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. In addition, the conversation data from the system is saved and can be reused, reducing the risk of copyright infringement.
[1568] System Components and Operation
[1569] 1. User Device
[1570] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1571] 2. Server
[1572] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[1573] Program processing flow
[1574] System startup and user authentication
[1575] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1576] Character selection and AI settings
[1577] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1578] Start a dialogue
[1579] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[1580] Response display and conversation continuation
[1581] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1582] Storage and secondary use of conversation data
[1583] The server stores the conversation data between the user and the generated AI in a database in real time as a log. Users can reuse their conversation data as needed, such as creating audio files, and the server converts the data for this purpose.
[1584] Exit and log out
[1585] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1586] Specific examples
[1587] For example, consider the process of a user logging in, selecting a particular famous author, and starting a dialogue. When the user types and sends a question such as, "What do you think of the latest novel?", the server passes the message to an AI model, which generates an appropriate response (e.g., "I've tried a particular subject. I'm curious to know how it resonates with you.") and sends it back to the device. The device displays this response to the user, and the dialogue continues. This series of processes allows the user to have a realistic dialogue experience.
[1588] The processing flow will be explained below.
[1589] Step 1:
[1590] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, user interface module) and connects to the database.
[1591] Step 2:
[1592] The user enters a username and password into the login page on the terminal.
[1593] Step 3:
[1594] The terminal transmits the login information entered by the user to the server.
[1595] Step 4:
[1596] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is taken to the main screen.
[1597] Step 5:
[1598] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[1599] Step 6:
[1600] The terminal transmits the information of the selected person to the server.
[1601] Step 7:
[1602] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking style to complete the initialization.
[1603] Step 8:
[1604] The user enters a text message and presses the send button to initiate a conversation.
[1605] Step 9:
[1606] The terminal transmits the input message to the server.
[1607] Step 10:
[1608] The server passes the received message to the AI model, which generates the optimal response to that message.
[1609] Step 11:
[1610] The server sends the generated response message to the terminal.
[1611] Step 12:
[1612] The terminal displays the response message received from the server to the user.
[1613] Step 13:
[1614] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1615] Step 14:
[1616] The server stores the interaction data between the user and the generated AI in a database in real time.
[1617] Step 15:
[1618] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[1619] Step 16:
[1620] The terminal sends a request to end the conversation to the server.
[1621] Step 17:
[1622] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[1623] Example 1
[1624] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1625] In recent years, dialogue systems using generative AI that mimic the personality and speaking style of a specific person have been attracting attention, but existing systems have not been able to provide realistic and personalized dialogue to reduce users' feelings of loneliness, limiting the experience.In addition, functions related to saving and secondary use of dialogue data are limited, presenting technical challenges for further enriching the user experience.
[1626] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1627] In this invention, the server includes means for initializing modules and connecting to a database when the system starts up, means for verifying user authentication information against the database and starting a user session if authentication is successful, means for allowing a user to select a specific person, acquiring conversation data related to that person, and loading and configuring a corresponding AI model, and means for storing the dialogue data in the database in real time and providing it in a reusable format as needed, thereby enabling users to smoothly engage in real, personalized dialogue with specific people and facilitating the storage and secondary use of the dialogue data.
[1628] A "user" refers to an entity that uses a specific terminal to access the dialogue system and engage in dialogue.
[1629] A "terminal" is a device used by a user to interact, and includes a PC, smartphone, tablet, etc.
[1630] A "server" is a computer system that plays a central role in the system, performing user authentication, managing interaction data, and operating the generation AI.
[1631] "Generative AI" refers to artificial intelligence models that mimic the personality and speaking style of a specific person to provide realistic dialogue.
[1632] "Personality" refers to a person's unique character traits and distinctive style of expression.
[1633] "Speech style" refers to a particular person's tone of voice and the way they use language.
[1634] A "prompt sentence" refers to a text message entered into the generation AI, based on which the AI generates a response.
[1635] A "user session" refers to session information for managing a series of interactions while a user is logged into a system.
[1636] A "database" is a system for storing and managing various data such as dialogue data and user information.
[1637] A "module" refers to a program component for realizing a specific function of a system.
[1638] "Initialization" refers to the process by which a system or module prepares to start operation.
[1639] "Logging out" refers to a user exiting the system and ending the user session.
[1640] "Interaction Data" refers to records of messages exchanged between the user and the generating AI.
[1641] "Secondary use" refers to converting the saved dialogue data into another format, such as an audio file, or using it for other purposes.
[1642] This section describes a specific embodiment of the present invention. This system allows users to select a specific person and provides realistic dialogue through a generative AI that mimics that person's personality and speaking style. Each of the following steps details the configuration and operation of the respective hardware and software.
[1643] System Configuration
[1644] The system mainly consists of the following components:
[1645] 1. User Device
[1646] A terminal is a device with which a user interacts, and includes a PC, smartphone, tablet, etc. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1647] 2. Server
[1648] The server plays a central role in the system, performing user authentication, dialogue data management, and operation management of the generative AI. The server is composed of multiple modules (e.g., a user authentication module, a data management module, an AI engine, etc.).
[1649] System Operation
[1650] The main system operations are explained below.
[1651] 1. System startup and user authentication
[1652] When the server starts the system, all necessary modules are initialized and connected to the database.
[1653] When a user enters login information on a terminal, the terminal sends that information to the server, which checks the received authentication information against a database and starts a user session if authentication is successful. If authentication fails, it returns an error message to the terminal.
[1654] 2. Character selection and AI settings
[1655] The user selects a specific person with whom they wish to interact on the platform, and the device sends the selected person's information to the server.
[1656] Based on the received person ID, the server retrieves the conversation data related to that person from the database, loads the corresponding AI model, and sets the selected person's personality and speaking style to the AI model, completing the initialization.
[1657] 3. Start the dialogue
[1658] The user inputs a text message to start a conversation and presses the send button. The terminal sends the input message to the server.
[1659] The server passes the received message to the AI model, which generates the optimal response to that message, which is then sent back to the device, where it is displayed to the user.
[1660] 4. Response display and conversation continuation
[1661] When the user enters and sends a new message to continue the conversation, the terminal again sends the message to the server.
[1662] The server again passes the message to the AI model and generates a response, and this process repeats until the interaction is over.
[1663] 5. Storage and secondary use of conversation data
[1664] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[1665] Users can reuse their dialogue data as audio files, etc., as needed, and the server converts the data accordingly. For example, to convert text data into audio data, a Text-to-Speech (TTS) engine is used.
[1666] 6. Exiting and Logging Out
[1667] When the user wants to end the conversation, he clicks the end button, and the terminal sends a request to end the conversation to the server.
[1668] The server ends the session, stores any necessary logs, and finally logs the user out and exits the system.
[1669] Specific examples
[1670] For example, consider the process of a user logging in and selecting a particular famous person (e.g., a famous scientist) to begin a dialogue. When the user types and submits a question such as, "What do you think about the latest scientific research?", the server passes the message to an AI model, which generates an appropriate response (e.g., "The latest research is very interesting, especially its original approach.") and sends it back to the device. The device then displays this response to the user and continues the dialogue.
[1671] Prompt Sentence Examples
[1672] For example, you might input the following prompts into a generative AI model:
[1673] "What do you think about the latest scientific research?"
[1674] The prompt allows the user to ask a question to the AI model as a specific person, with the expectation that it will generate a response that mimics the way that person speaks and thinks.
[1675] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1676] Step 1: System startup and module initialization
[1677] When the server starts the system, it initializes all necessary modules (user authentication module, data management module, AI engine, etc.) and connects to the database.
[1678] Input: System startup trigger.
[1679] Processing: Module initialization, database connection establishment.
[1680] Output: System initialization complete and database connection status.
[1681] Step 2: User authentication
[1682] The user enters login information on the device and clicks the login button.
[1683] The terminal sends the entered login information (user name, password) to the server.
[1684] The server checks the received authentication information against the user information in its database, and if authentication is successful, starts a user session.
[1685] Input: Username, Password.
[1686] Processing: Database verification, session start.
[1687] Output: Session ID, authentication success or failure result.
[1688] Step 3: Selecting a person and setting AI
[1689] The user selects the specific person with whom they wish to interact on the platform.
[1690] The device sends the selected person's information (person ID and name) to the server.
[1691] Based on the received person information, the server retrieves conversation data related to that person from the database and loads the corresponding AI model.
[1692] The server sets the AI model with the selected person's personality and speaking style, completing initialization.
[1693] Input: Person ID, Name.
[1694] Processing: Acquire conversation data, load AI model, and set personality.
[1695] Output: The initialized AI model.
[1696] Step 4: Start a conversation
[1697] The user enters a text message to initiate a conversation and presses the send button.
[1698] The terminal transmits the input message to the server.
[1699] The server passes the received message to the AI model, which generates the optimal response to that message.
[1700] The generated response is sent back from the server to the terminal, which displays the response to the user.
[1701] Input: User message.
[1702] Processing: Parse the message, feed it into the AI model, and generate a response.
[1703] Output: The generated response message.
[1704] Step 5: Keep the conversation going
[1705] The user enters and sends a new message to continue the conversation.
[1706] The terminal sends a new message to the server.
[1707] The server again passes the message to the AI model, which generates a response.
[1708] This process is repeated until the conversation is over.
[1709] Input: New user message.
[1710] Processing: Parse the message, feed it into the AI model, and generate a response.
[1711] Output: The generated response message.
[1712] Step 6: Saving and Reusing Conversation Data
[1713] The server stores the conversation data between the user and the generated AI as a log in a database in real time.
[1714] The user can reuse the dialogue data as a voice file or the like.
[1715] The server performs the data conversion for this purpose, for example by using a Text-to-Speech (TTS) engine.
[1716] Input: Interaction data.
[1717] Processing: Data storage, data conversion (e.g. text to speech).
[1718] Output: Saved interaction logs, transformed data.
[1719] Step 7: End the conversation and log out
[1720] If the user wishes to end the interaction, he clicks on the End button.
[1721] The terminal sends a request to end the conversation to the server.
[1722] The server terminates the session and saves any necessary logs.
[1723] Finally, the server logs the user out and causes him to leave the system.
[1724] Input: A conversation termination request.
[1725] Processing: Session termination, log saving, logout processing.
[1726] Output: Finished sessions, saved interaction logs, logout status.
[1727] (Application example 1)
[1728] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1729] Conventional dialogue systems allow users to have a realistic dialogue experience through dialogue with a generative AI that mimics the personality and speaking style of a specific person. However, there is a need to provide a richer dialogue experience by adding entertainment and information functions. Furthermore, there is a lack of functionality that allows users to save their dialogue history and use it for later playback or learning. Technology is needed to solve these issues.
[1730] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1731] In this invention, the server includes means for retrieving conversation data related to a selected person from a database and loading and configuring a corresponding generative AI model, means for verifying user authentication information and starting a user session if login is successful, and means for adjusting and optimizing the generative AI model for providing entertainment and information. This allows a user to enjoy entertainment and information in real time through conversations with a specific person, and saves the conversation history for later playback or learning.
[1732] A "user" is an entity that uses the system to have a conversation with a specific person.
[1733] A "specific person" is someone the user selects and a dialogue is provided through a generative AI that mimics that person's personality and speaking style.
[1734] "Personality" refers to the characteristics, traits, and speaking style that are unique to a particular person.
[1735] "Generative AI" is artificial intelligence that mimics the personality and speaking style of a specific person and generates realistic conversations with the user.
[1736] A "text message" is the text information entered by the user, and is the data that the generation AI uses to generate a response.
[1737] A "server" is a device that plays a central role in the system, performing user authentication, managing dialogue data, and operating the generation AI.
[1738] A "response" is a reply message that the generation AI generates in response to a user's text message.
[1739] A "terminal" is a device used by a user to interact, and includes a smartphone, a personal computer, etc.
[1740] "Dialogue data" is a record of the conversation between the user and the generating AI.
[1741] "Secondary use" refers to reusing saved dialogue data in a different format or for a different purpose.
[1742] "Logging out" means ending an interactive session and exiting the system by the user.
[1743] "Conversation data" is dialogue information related to a particular person.
[1744] A "database" is a system that stores and manages information such as conversation data.
[1745] An "AI model" is an algorithm or mathematical model that serves as the basis for generative AI that mimics the personality and speaking style of a specific person.
[1746] "Entertainment" refers to interactive content and functions that provide enjoyment to users.
[1747] "Information" refers to knowledge and data that users acquire through dialogue.
[1748] "Playback" refers to using the saved dialogue history to check it later.
[1749] "Logging in" is the process by which a user enters and confirms authentication information in order to access and use a system.
[1750] "Authentication information" refers to personal information required to access a system, such as a user ID and password.
[1751] A "user session" is a series of operations performed by a user from the time the user logs in until the time the user logs out.
[1752] An "error message" is a warning or guidance message that is displayed to the user when login fails.
[1753] "Tuning" is the process of configuring and optimizing a generative AI model.
[1754] "Optimization" refers to adjusting the generative AI model to improve the user's interaction experience.
[1755] As an embodiment of the present invention, a dialogue platform system using a generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue is described below. This system is realized when a user selects a specific person and sends a message to the generative AI.
[1756] Overall system overview
[1757] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through a generative AI. This system is designed to help users reduce their sense of loneliness, and is particularly useful for parents raising children and lonely elderly people. In addition, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement.
[1758] System Components and Operation
[1759] 1. User Device
[1760] The terminal is a device with which the user interacts, and can be a smartphone, smart glasses, or a head-mounted display. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1761] 2. Server
[1762] The server plays a central role in the system, performing user authentication, dialogue data management, and generating AI operations. The server is composed of multiple modules.
[1763] Program processing flow
[1764] 1. System startup and user authentication
[1765] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1766] 2. Character selection and AI settings
[1767] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1768] 3. Start the dialogue
[1769] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to the AI model, which generates the optimal response to the message. The generated response is sent back to the device, which then displays it to the user.
[1770] 4. Response display and conversation continuation
[1771] When the user types and sends a new message to continue the conversation, the device again sends the message to the server, which again passes the message to the AI model and generates a response. This process repeats until the conversation ends.
[1772] 5. Storage and secondary use of conversation data
[1773] The server stores the dialogue data between the user and the generated AI in a database in real time as a log. The user can reuse the dialogue data as audio files, etc., as needed, and the server converts the data for this purpose.
[1774] 6. Exiting and Logging Out
[1775] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1776] Hardware and software used
[1777] Hardware: Smartphones, smart glasses, head-mounted displays, servers
[1778] Software: Web applications, native applications, generative AI models, databases (MySQL, PostgreSQL), cloud services (AWS, Google Cloud)
[1779] Examples of specific examples and prompts
[1780] For example, consider a scenario in which a user logs into a "virtual conversation hall" and begins a conversation with a famous author. When the user types and sends a question such as "What do you think about the latest novel?", the server passes the message to an AI model, which generates an appropriate response. The generated response (e.g., "I've tried a particular theme. I'm curious to see how it resonates with you.") is sent back to the device and displayed to the user. In this way, the user can have a realistic conversation experience.
[1781] Prompt Sentence Examples
[1782] User: What do you think about the latest novel?
[1783] Server: Sends messages to the AI model and generates responses such as:
[1784] AI Model: "I've tried a particular theme and I'm curious to see how it resonates."
[1785] In this way, the present invention provides users with a realistic interaction experience with a specific person, enabling them to enjoy entertainment and information.
[1786] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1787] Step 1:
[1788] System startup and user authentication:
[1789] The server starts the system, initializes all necessary modules, and connects to the database. The user enters login information (username, password) on the terminal. The terminal sends the entered information to the server. The server compares the received login information with the database, and if authentication is successful, starts the user session. At this time, it obtains the user's authentication information from the database after hashing it and compares it with the entered information. If authentication fails, an error message is returned to the user's terminal. As output, if successful, the user session will start, and if unsuccessful, an error message will be displayed.
[1790] Step 2:
[1791] Character Selection and AI Settings:
[1792] The user selects a specific person with whom they wish to interact on the platform. The device sends the selected person's information to the server. The server uses this information to retrieve conversation data related to the selected person from a database. The server then loads the corresponding generative AI model, sets the selected person's personality and speaking style in the AI model, and completes initialization. The configured generative AI model is prepared as the output.
[1793] Step 3:
[1794] Start the conversation:
[1795] The user enters a text message to start a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to a generative AI model, which generates an optimal response to that message. In this case, the generative AI model (e.g., OpenAI GPT-3, GPT-4) takes text data as input and runs an algorithm to generate an appropriate response. The generated response is sent back from the server to the device, and the device displays it to the user. As an output, a response message is generated that is displayed on the user's device.
[1796] Step 4:
[1797] Response display and conversation continuation:
[1798] The user types and sends a new text message to continue the dialogue. The device again sends the message to the server. The server again passes the message to the generative AI model, which generates a new response. This process repeats until the dialogue ends. Each message exchange involves data processing and the generative AI generating a response based on the new user input. As an output, the continuing dialogue message is displayed on the user's device.
[1799] Step 5:
[1800] Conversation data storage and secondary use:
[1801] The server stores the dialogue data between the user and the generated AI in a database in real time. The stored dialogue data can be converted into an audio file or other format as needed for secondary use. The server uses a data conversion algorithm to convert text data into an audio file. As output, the stored dialogue data and data in a format that can be used for secondary use are generated.
[1802] Step 6:
[1803] Exit and log out:
[1804] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server. The server ends the session and saves the necessary logs. Finally, the user is logged out and exits the system. As an output, the user is safely logged out of the system.
[1805] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1806] As an embodiment of the present invention, a system is described below that combines an emotion engine with a dialogue platform using generative AI that mimics the personality and speaking style of a specific person and provides realistic dialogue. This system is realized when a user selects a specific person and sends a message to the generative AI, which recognizes the user's emotions and generates and provides a response based on them.
[1807] Overall system overview
[1808] In this system, the user initiates a conversation with a specific person using a device, and the server provides a realistic conversation with the user through generative AI. This system is designed to help reduce the user's sense of loneliness, and is particularly beneficial for mothers raising children and lonely elderly people. Furthermore, the system's conversation data is saved and can be reused, reducing the risk of copyright infringement. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and adjusts responses accordingly.
[1809] System Components and Operation
[1810] 1. User Device
[1811] A terminal is a device with which a user interacts, and can be a PC, smartphone, or tablet. The terminal is connected to the Internet and accesses the server through the system's web application or native application.
[1812] 2. Server
[1813] The server plays a central role in the system, performing user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[1814] 3. Generation AI
[1815] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue. The server sets the necessary parameters and generates a response based on the user's message.
[1816] 4. Emotion Engine
[1817] The emotion engine recognizes emotions from the text messages entered by the user and adjusts the generative AI's response based on those emotions. Emotion data is also stored along with the dialogue data and can be used for analysis and feedback as needed.
[1818] Program processing flow
[1819] System startup and user authentication
[1820] The server initializes all necessary modules at system startup and connects to the database. When a user enters login information on a terminal, the terminal sends that information to the server. The server compares the received authentication information with the database, and if authentication is successful, it starts a user session. If authentication fails, it returns an error message to the user terminal.
[1821] Character selection and AI settings
[1822] The user selects a specific person they wish to interact with on the platform. The device then sends the selected person's information to the server. The server then retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The server then sets the selected person's personality and speaking style in the AI model, completing the initialization.
[1823] Start a dialogue
[1824] The user enters a text message to initiate a dialogue and presses the send button. The device sends the entered message to the server. The server passes the received message to an emotion engine to recognize the user's emotion. The recognized emotion data is passed to an AI model, which generates a tailored response based on the emotion. The generated response is sent back to the device, which then displays it to the user.
[1825] Response display and conversation continuation
[1826] When the user enters and sends a new message to continue the conversation, the device again sends the message to the server, which then passes it to the emotion engine, which recognizes the user's emotion and generates an appropriate response. This process is repeated until the conversation ends.
[1827] Storage and secondary use of conversation data
[1828] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time. Users can reuse their conversation data as audio files, etc., as needed, and the server converts the data for this purpose. The server also analyzes the emotional data and uses it as feedback to improve the quality of the conversation.
[1829] Exit and log out
[1830] When the user wants to end the conversation, he clicks the end button. The terminal sends a request to end the conversation to the server, which then ends the session and saves the necessary logs. Finally, the user is logged out and exits the system.
[1831] Specific examples
[1832] For example, when a user logs in and selects a particular famous author to begin a dialogue, the user can type in a question like, "What do you think about the latest novel?" and send it. The server then passes the message to the emotion engine, which recognizes the user's emotion (e.g., excitement, curiosity). Based on the emotion data, the AI model generates a response (e.g., "I've tried a particular topic. I'm curious to see how it resonates with you.") and sends it back to the device. The device then displays this response to the user, continuing the dialogue. This process allows the user to have a more personal and realistic dialogue experience.
[1833] The processing flow will be explained below.
[1834] Step 1:
[1835] When the system starts up, the server initializes all necessary modules (database connection module, AI model module, emotion engine module, and user interface module) and connects to the database.
[1836] Step 2:
[1837] The user enters a username and password into the login page on the terminal.
[1838] Step 3:
[1839] The terminal transmits the login information entered by the user to the server.
[1840] Step 4:
[1841] The server verifies the received login information against the database and authenticates the user. If authentication is successful, the user session begins and the user is redirected to the main screen. If authentication fails, an error message is returned.
[1842] Step 5:
[1843] From the main screen, the user selects a particular person (e.g., a famous author or fictional character) with whom they wish to interact.
[1844] Step 6:
[1845] The terminal transmits the information of the selected person to the server.
[1846] Step 7:
[1847] The server retrieves the conversation data related to the selected person from the database, loads the corresponding AI model, and sets the AI model with the person's personality and speaking characteristics to complete initialization.
[1848] Step 8:
[1849] The user enters a text message and presses the send button to initiate a conversation.
[1850] Step 9:
[1851] The terminal transmits the input message to the server.
[1852] Step 10:
[1853] The server passes the received message to the emotion engine, which analyzes the text and recognizes the user's emotion from the message.
[1854] Step 11:
[1855] The server passes the recognized user emotion data to the AI model, which uses it to generate an appropriate response.
[1856] Step 12:
[1857] The server generates a response message adjusted based on the emotion data and transmits it to the terminal.
[1858] Step 13:
[1859] The terminal displays the response message received from the server to the user.
[1860] Step 14:
[1861] When the user enters and sends a new message to continue the conversation, the terminal sends the message to the server again.
[1862] Step 15:
[1863] The server again passes the message to the emotion engine, which recognizes the emotion and generates a response based on the recognized emotion data. This process is repeated until the interaction is completed.
[1864] Step 16:
[1865] The server stores the conversation data and emotion data between the user and the generated AI in a database in real time.
[1866] Step 17:
[1867] When the user has finished interacting and wishes to log out of the system, he clicks on the Exit button.
[1868] Step 18:
[1869] The terminal sends a request to end the conversation to the server.
[1870] Step 19:
[1871] The server ends the interactive session, saves any necessary logs, and finally logs the user out and causes them to leave the system.
[1872] Example 2
[1873] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1874] In modern society, many users experience feelings of loneliness, a major problem especially for mothers raising children and lonely elderly people. Conventional dialogue systems struggle to fully recognize users' emotions and respond appropriately, preventing users from engaging in satisfying dialogue. Furthermore, data integrity and security are not guaranteed when dialogue data is stored or reused. A system that can address these issues is needed.
[1875] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotions using an emotion engine based on a message received by the server and generating an appropriate response, means for returning the generated response to the terminal, and means for the server to store dialogue data and emotion data between the user and the generation AI and provide them in a format that can be reused. This allows the user to experience a realistic dialogue and obtain an appropriate response based on the emotion. Furthermore, since the stored dialogue data and emotion data can be reused, it is possible to improve the quality of the dialogue and increase user satisfaction.
[1876] "User" refers to a person who uses the system to interact with a specific person.
[1877] "Terminal" refers to a device such as a PC, smartphone, or tablet that a user uses to interact.
[1878] A "server" refers to a computer that plays a central role in the system and performs user authentication, dialogue data management, generative AI operation, emotion recognition, etc.
[1879] "Generative AI" refers to an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[1880] An "emotion engine" refers to a program that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[1881] "Dialogue data" refers to data that records the content of the dialogue between the user and the generating AI.
[1882] "Secondary use" refers to converting saved dialogue data and emotion data into audio files, etc., and reusing them for analysis or other purposes.
[1883] "Initialization" refers to the process of setting up a system or program so that it can operate normally.
[1884] "Personality" refers to the characteristics and traits that a particular person possesses.
[1885] "Logging in" refers to the process by which a user enters authentication information to access a system and is authenticated by the system.
[1886] A "session" refers to a series of interactions with the system from the time a user logs in until the time they log out.
[1887] This invention is a dialogue platform system that uses generative AI to mimic the personality and speaking style of a specific person and provide realistic dialogue. When a user selects a specific person and sends a message to the generative AI, the AI recognizes the user's emotions and generates and provides a response based on those emotions. This system is particularly useful for users who feel lonely, such as mothers raising children or lonely elderly people.
[1888] System Components
[1889] 1. User Device
[1890] Terminals include devices connected to the Internet, such as PCs, smartphones, tablets, etc. Users interact with these terminals.
[1891] The terminal accesses the server through a web application or a native application.
[1892] 2. Server
[1893] The server is the core of the system, and is responsible for user authentication, dialogue data management, generative AI operation, emotion recognition, etc. The server is composed of multiple modules.
[1894] The server sets the necessary parameters and generates a response according to the user's message.
[1895] 3. Generation AI
[1896] Generative AI is an artificial intelligence program that mimics the personality and speaking style of a specific person to generate realistic dialogue.
[1897] The generative AI creates a realistic interaction experience for users based on the characteristics of the person they select.
[1898] 4. Emotion Engine
[1899] The emotion engine is responsible for recognizing emotions from the text messages entered by the user and adjusting the generative AI's response based on those emotions.
[1900] Emotional data is stored along with the dialogue data and can be used for analysis and feedback as needed.
[1901] System Operation
[1902] In this system, the user initiates a conversation using a device, and the server provides a realistic conversation via a generative AI. The overall flow of the system is shown below.
[1903] 1. User authentication: The user enters login information on the terminal, which then sends it to the server. The server checks the authentication information against a database and starts the session if authentication is successful. If authentication fails, it returns an error message.
[1904] 2. Person selection: The user selects the desired conversation partner, and the device sends this information to the server. The server retrieves the conversation data related to the selected person from the database and loads the corresponding AI model. The AI model initialization is complete.
[1905] 3. Interaction: When the user inputs a message, the device sends it to the server. The server passes the message to the emotion engine to recognize the user's emotion. It generates a response tailored based on the emotion and sends it back to the device. The device displays the response to the user.
[1906] 4. Continuous Dialogue: This process repeats as long as the user continues to input new messages. For each message, the appropriate emotion is recognized and a response is generated.
[1907] 5. Data storage and secondary use: The server stores the conversation data and emotion data between the user and the generated AI in a database. The user can reuse the conversation data as audio files, and the server converts the data for this purpose. The server also analyzes the emotion data and uses it as feedback to improve the quality of the conversation.
[1908] 6. Termination and logout: When the user finishes the conversation, the terminal sends a request to terminate the conversation to the server, which then terminates the session, saves the necessary logs, and logs out the user.
[1909] Specific examples of operation
[1910] For example, a user logs into the system, selects a particular well-known author, and begins a dialogue by typing and sending a question such as, "What do you think of his latest novel?" The server passes this message to the emotion engine, which recognizes the user's emotion (e.g., excitement or curiosity). Based on this emotion data, the generative AI generates a response such as, "I've tried a particular theme. I'm curious to see how it resonates with you," and sends it back to the device. The device displays this response to the user, allowing the user to continue the dialogue.
[1911] Prompt Sentence Examples
[1912] "What is your opinion on the latest science and technology?"
[1913] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1914] Step 1:
[1915] System startup
[1916] The server initializes all modules in the system, securing the necessary resources and establishing connections to the database.
[1917] Input: System startup instructions
[1918] Data processing: Module initialization, resource allocation
[1919] Output: Initialization complete, connection established
[1920] Step 2:
[1921] User Authentication
[1922] The user enters login information (user name, password) on the terminal and presses the send button.
[1923] The terminal transmits the entered authentication information to the server.
[1924] The server receives the input information and checks it against a database.
[1925] Input: Authentication information (username, password)
[1926] Data processing: Matching with database
[1927] Output: Authentication success or failure message
[1928] If authentication is successful, the server starts a user session and generates a corresponding session ID, otherwise it returns an error message to the terminal.
[1929] Step 3:
[1930] Character Selection
[1931] Users select the specific person on the platform they want to interact with.
[1932] The terminal transmits the information of the selected person to the server.
[1933] Based on the received information, the server retrieves relevant conversation data from the database and loads the corresponding AI model.
[1934] Input: Selected person's information
[1935] Data processing: Acquire conversation data, load AI model
[1936] Output: AI model initialization complete
[1937] Step 4:
[1938] Start a dialogue
[1939] The user enters a message and presses the send button.
[1940] The terminal transmits the input message to the server.
[1941] The server passes the received message to the emotion engine to recognize the user's emotion.
[1942] Input: User message
[1943] Data processing: Emotion recognition
[1944] Output: Emotion data
[1945] The generation AI generates a response based on the recognized emotion data, and the server sends the generated response back to the device.
[1946] Step 5:
[1947] Response Display
[1948] The terminal displays the response received from the server to the user.
[1949] Input: Generator AI response
[1950] Data processing: Displaying responses
[1951] Output: The response displayed on the user's screen
[1952] Step 6:
[1953] Ongoing dialogue
[1954] The user enters a new message and repeats steps 4 and 5 above.
[1955] Input: New user message
[1956] Data processing: emotion recognition, response generation
[1957] Output: Shows the new response
[1958] Step 7:
[1959] Saving conversation data
[1960] The server stores the conversation data and emotional data between the user and the generated AI in a database in real time.
[1961] Input: Dialogue data, emotion data
[1962] Data processing: Saving to database
[1963] Output: Saved interaction data
[1964] Step 8:
[1965] Secondary use
[1966] The user can reuse the saved dialogue data as a voice file or the like.
[1967] The server converts the interaction data and provides it in the specified format.
[1968] Input: Saved interaction data
[1969] Data processing: Data conversion
[1970] Output: Converted data format
[1971] Step 9:
[1972] Exit and log out
[1973] If the user wishes to end the dialogue, he or she presses the end button.
[1974] The terminal sends a request to end the conversation to the server.
[1975] The server ends the session, saves any necessary logs, and logs the user out.
[1976] Input: Conversation termination request
[1977] Data processing: End session, save log
[1978] Output: Logout completed
[1979] (Application example 2)
[1980] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1981] Current dialogue systems for factory robots have difficulty achieving natural communication with humans and are unable to provide sufficient support to workers. As a result, it is difficult to improve work efficiency and safety, and there is a risk of worker fatigue and stress increasing. In addition, the robot lacks the ability to recognize workers' emotions and respond appropriately based on them, which means that there is an issue of not being able to reduce workers' psychological burden or provide sufficient support for their work.
[1982] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user's emotion and generating a response based on that emotion, means for the server to recognize the user's emotion and adjust the response based on the emotion data, and means for the user to receive work instructions through dialogue with the robot. This not only enables natural communication when factory workers interact with the robot, but also enables appropriate responses and support that take into account the worker's emotion and fatigue. As a result, it is expected that the psychological burden on workers will be reduced and work efficiency and safety will be improved.
[1983] "User" refers to a human being who uses the system to engage in dialogue using generative AI.
[1984] A "specific person" refers to a person selected by the user whose personality and speaking style the generated AI will imitate and set up to converse with.
[1985] "Generative AI" refers to artificial intelligence programs that are designed to mimic the personality and speaking style of a specific person.
[1986] "Server" refers to the central computer system that manages the storage of interaction data, user authentication, emotion recognition, and the operation of generative AI.
[1987] An "emotion engine" refers to a module that recognizes emotions from messages entered by users and generates responses based on those emotions.
[1988] A "terminal" is a device with which a user interacts, such as a smartphone, tablet, or computer.
[1989] "Dialogue data" refers to information including messages and emotional data exchanged between the user and the generating AI.
[1990] "Secondary use" refers to reusing saved dialogue data for other purposes.
[1991] "AI model" refers to an artificial intelligence module that includes the configuration information and algorithms used by the generative AI.
[1992] "Emotion data" refers to data that recognizes a user's emotions and includes information about them.
[1993] A "robot" refers to a mechanical device used to interact with workers in factories and other places.
[1994] "Work instructions" refer to instructions regarding specific work content that a user receives through interaction with a robot.
[1995] "Login information" refers to information used for user authentication, and primarily refers to user IDs and passwords.
[1996] This invention is a system that uses a support robot to interact with factory workers. The system is designed to allow users to experience realistic interactions through a generative AI that mimics the personality and speaking style of a specific person. The specific configuration and operation for implementing this system are described below.
[1997] Overall system configuration
[1998] The system mainly consists of the following components:
[1999] 1. User Device:
[2000] The devices used are smartphones, tablets, and computers. These devices connect to the internet and communicate with the server.
[2001] 2. Server:
[2002] The Server is the central computer system that processes user input messages and generates responses using generative AI. The Server contains the following main modules:
[2003] User Authentication Module
[2004] Generative AI Module
[2005] Emotion Engine Module
[2006] Database Management Module
[2007] 3. Generation AI:
[2008] Generative AI is an artificial intelligence program that generates realistic dialogue by imitating the personality and speaking style of a specific person, and is primarily implemented using the GPT-2 model.
[2009] 4. Emotion Engine:
[2010] The emotion engine is a module that recognizes emotions from text messages entered by users and adjusts the generative AI's responses based on those emotions.
[2011] Operation overview
[2012] The system operation is outlined below.
[2013] 1. User authentication:
[2014] The user enters login information on the terminal and sends it to the server, which verifies the information and, if authentication is successful, starts the user session.
[2015] 2. Selecting a specific person:
[2016] The user selects a specific person with whom they wish to interact, and the device transmits the selection to the server, which loads and configures the corresponding generative AI model.
[2017] 3. Start a conversation:
[2018] To initiate a dialogue, a user inputs and sends a text message, and the server passes the message to an emotion engine, which recognizes the user's emotion, generates an appropriate response, and sends it back to the device.
[2019] 4. Continuing the conversation and storing data:
[2020] As the conversation continues, new messages from the user are processed in the same way. All conversation data is stored on the server in real time and made available in a reusable format.
[2021] Usage example
[2022] For example, imagine a factory worker logs into the system and selects a specific support representative to begin a conversation. If the user types, "Please tell me how to operate this machine," the following prompt sentence is generated:
[2023] Example prompt sentence:
[2024] The following is a conversation with a supportive robot. The robot is very helpful and supportive.
[2025] Human: Please tell me how to operate this machine.
[2026] Robot (supportive):
[2027] The server uses a generative AI model to generate a response, such as, "First, turn on the machine, then follow the instructions on the control panel."
[2028] The system allows factory workers to receive specific work instructions through a realistic interactive experience, which is expected to improve efficiency and safety.
[2029] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2030] Step 1:
[2031] User authentication: The user enters login information (user ID and password) on the terminal and sends it to the server. The server checks the information against a database, and if authentication is successful, a user session begins. If authentication is successful, the server generates a session ID and returns it to the terminal. Conversely, if authentication fails, the server returns an error message.
[2032] Input: User ID, Password
[2033] Data processing: Database matching of authentication information
[2034] Output: Session ID (if successful), Error message (if unsuccessful)
[2035] Step 2:
[2036] Selecting a specific person: The user selects a specific person with whom they wish to converse. The device sends the selected person's information to the server. The server retrieves the conversation data related to that person from the database, loads and configures the generative AI model, and then configures the model with the specific person's personality and speaking style.
[2037] Input: Selected person's information
[2038] Data processing: Acquiring a database of conversation data, loading and configuring the generative AI model
[2039] Output: The configured generative AI model
[2040] Step 3:
[2041] Dialogue start: The user enters a text message to start the dialogue and sends it from the device to the server. The server passes the received message to the emotion engine, which recognizes the user's emotion. The recognized emotion data is passed to a generative AI model, which generates an appropriate response based on the emotion. The generated response is then sent back to the device and displayed to the user.
[2042] Input: User's text message
[2043] Data processing: Emotion recognition and response generation using generative AI
[2044] Output: Response message
[2045] Step 4:
[2046] Continuing the dialogue and saving data: Every time the user inputs a new message to continue the dialogue, the device sends the message to the server. Emotion recognition is performed in the same way, and a response is generated based on the emotional data. This series of processes is then repeated. Furthermore, the server saves all dialogue data in a database.
[2047] Input: The user's new text message
[2048] Data processing: Emotion recognition, response generation using generative AI, and dialogue data storage
[2049] Output: Response message, saved interaction data
[2050] Step 5:
[2051] Ending the conversation and logging out: When the user wants to end the conversation, he clicks the end button and the terminal sends a request to end the conversation to the server. The server ends the conversation session and saves the log information. After that, the user is logged out and exits the system.
[2052] Input: Conversation termination request
[2053] Data processing: End session, save log information
[2054] Output: Logged out of the system
[2055] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2056] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2057] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2058] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2059] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2060] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2061] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2062] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2063] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2064] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2065] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2066] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2067] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2068] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2069] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2070] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2071] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2072] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2073] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2074] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2075] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2076] The following is further disclosed regarding the above embodiment.
[2077] (Claim 1)
[2078] A way for users to select a specific person and experience a realistic conversation through generative AI that mimics that person's personality and speaking style.
[2079] means for receiving a text message entered by a user and transmitting the text message to a server;
[2080] a means for the server to generate an appropriate response based on the received message and return it to the terminal;
[2081] A means for the server to store the dialogue data between the user and the AI and provide it in a format that can be reused as needed;
[2082] a means for the user to end the interaction and log out;
[2083] A system including:
[2084] (Claim 2)
[2085] means for the server to retrieve conversational data associated with the selected person from the database and load and configure a corresponding AI model;
[2086] A means for the server to set the personality and speaking characteristics of the target person in the AI model and complete initialization;
[2087] The system of claim 1 further comprising:
[2088] (Claim 3)
[2089] a means by which the server validates the user's credentials and initiates a user session if the login is successful;
[2090] A means to return an error message when login fails,
[2091] The system of claim 1 further comprising:
[2092] "Example 1"
[2093] (Claim 1)
[2094] A way for users to select a specific person and experience a realistic conversation through generative AI that mimics that person's personality and speaking style.
[2095] means for receiving a text message entered by a user and transmitting the text message to a server;
[2096] a means for the server to generate an appropriate response based on the received message and return it to the terminal;
[2097] A means for the server to store the dialogue data between the user and the AI and provide it in a format that can be reused as needed;
[2098] A means for the server to initialize the module and connect to the database at system startup;
[2099] a means for the user to end the interaction and log out;
[2100] A system including:
[2101] (Claim 2)
[2102] means for the server to retrieve conversational data associated with the selected person from the database and load and configure a corresponding AI model;
[2103] A means for the server to set the personality and speaking characteristics of the target person in the AI model and complete initialization;
[2104] ...
Claims
1. A way for users to select a specific person and experience a realistic conversation through generative AI that mimics that person's personality and speaking style. means for receiving a text message entered by a user and transmitting the text message to a server; a means for the server to generate an appropriate response based on the received message and return it to the terminal; A means for the server to store the dialogue data between the user and the AI and provide it in a format that can be reused as needed; a means for the user to end the interaction and log out; A system including:
2. means for the server to retrieve conversational data associated with the selected person from the database and load and configure a corresponding AI model; A means for the server to set the personality and speaking characteristics of the target person in the AI model and complete initialization; The system of claim 1 further comprising:
3. a means by which the server validates the user's credentials and initiates a user session if the login is successful; A means to return an error message when login fails, The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A