System

The system addresses the limitations of conventional generative AI by tailoring AI personalities to user needs through morphological analysis and interaction log analysis, enhancing user engagement and satisfaction.

JP2026025482APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128291
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Conventional generative AI systems provide mechanical and uninteresting interactions, failing to meet individual user needs and resulting in user dissatisfaction.

Method used

A system that includes morphological analysis of user input to extract personality traits, generates an AI personality tailored to these traits, initiates interactive sessions, saves interaction logs, and analyzes them to improve response patterns, enabling real-time, two-way communication.

Benefits of technology

The system provides immersive and satisfying communication by generating AI personalities that meet individual user needs, ensuring personalized and engaging interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025482000001_ABST
    Figure 2026025482000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a user input; means for morphologically analyzing text data inputted by the user; means for extracting a personality characteristic from the analyzed text data; means for generating an AI personality based on the extracted personality characteristic; means for presenting the generated AI personality to the user; means for starting a dialogue session with the AI personality selected by the user; means for storing exchanges during the dialogue session; and means for analyzing the stored dialogue log and improving a response pattern.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Chatting with conventional generative AI was mechanical and uninteresting, leading users to quickly become bored. Furthermore, they could only provide general advice that applied to everyone, making it difficult to meet individual needs. As a result, conventional generative AI systems failed to satisfy users. The present invention aims to solve these problems by providing immersive communication for users, offering useful advice tailored to their individual needs and providing a sense of satisfaction. [Means for solving the problem]

[0005] The present invention solves these problems by including a means for accepting user input, a means for morphologically analyzing the text data entered by the user, a means for extracting personality traits from the analyzed text data, a means for generating an AI personality based on the extracted personality traits, a means for presenting the generated AI personality to the user, a means for starting an interaction session with the AI ​​personality selected by the user, a means for saving exchanges during the interaction session, and a means for analyzing the saved interaction log and improving response patterns. This allows users to engage in real-time, two-way communication with interlocutors that meet their specific needs, resulting in a fulfilling and satisfying experience.

[0006] "User" refers to an individual who interacts with the System.

[0007] "Input" refers to the act of a user providing text data to a system or the data itself.

[0008] "Text data" refers to a collection of sentences or words that a user inputs into a system.

[0009] "Morphological analysis" refers to a method of analyzing the meaning and structure of words and phrases that make up text data.

[0010] "Personality traits" refer to features that represent the personality and behavioral patterns of the interlocutor, extracted from morphological analysis.

[0011] "AI personality" refers to a virtual interlocutor generated based on extracted personality traits.

[0012] "Presenting" refers to the act of a terminal displaying information to a user.

[0013] "Interaction session" refers to a period of continuous communication between a user and a selected AI personality.

[0014] "Interaction Log" means data that records all interactions during an interaction session.

[0015] "Analysis" refers to the act of examining stored dialogue logs in detail to find meaning and patterns.

[0016] "Response pattern" refers to the format and content of the AI ​​personality's responses to the user during an interactive session.

[0017] "Training data" refers to historical interaction logs and other data sets used to refine the response patterns of an AI model. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user. Below, as an embodiment of the present invention, the processing of the system's program is explained in detail in natural language.

[0040] Basic system configuration

[0041] The system consists of the following elements:

[0042] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0043] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[0044] Database: A database that stores interaction logs and training data.

[0045] System processing flow

[0046] 1. The user launches the application

[0047] The user taps the application icon on their smartphone to launch it, and the device displays the application's home screen.

[0048] 2. Accept user input

[0049] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, the user might input, "Someone who is quiet and gives logical advice."

[0050] 3. Sending input data

[0051] The terminal transmits the input text data to the server.

[0052] 4. Morphological analysis and feature extraction

[0053] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." This analysis is performed using natural language processing technology.

[0054] 5. AI personality generation

[0055] The server generates an appropriate AI personality based on the extracted keywords, for example, selecting and customizing an AI model with a response pattern that matches a specific personality trait (e.g., quiet and logical).

[0056] 6. Present interlocutor options

[0057] The server creates a list of multiple AI personalities and sends it to the device. The device then displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The user then selects the desired interlocutor.

[0058] 7. Starting an interactive session

[0059] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0060] 8. Real-time two-way communication

[0061] The device sends the questions or inquiries entered by the user to the server, which then uses its AI personality to generate an appropriate response and sends it back to the device, which then displays it to the user. This process is repeated, enabling real-time two-way communication.

[0062] 9. Saving conversation logs

[0063] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0064] 10. Analyzing conversation logs and improving response patterns

[0065] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0066] Specific examples

[0067] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what is bothering her at work today," the server generates Ken's response and provides appropriate advice. This conversation log is saved on the server and used for future conversations.

[0068] As described above, the present invention allows users to communicate in real time with the interlocutor that best suits them at any time, providing an immersive and satisfying experience.

[0069] The processing flow will be explained below.

[0070] Step 1:

[0071] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0072] Step 2:

[0073] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0074] Step 3:

[0075] The terminal transmits the input text data to the server.

[0076] Step 4:

[0077] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0078] Step 5:

[0079] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0080] Step 6:

[0081] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[0082] Step 7:

[0083] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[0084] Step 8:

[0085] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0086] Step 9:

[0087] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[0088] Step 10:

[0089] The server generates a response from the AI ​​personality "Ken" and sends it back to the device. For example, the server generates logical advice as text.

[0090] Step 11:

[0091] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[0092] Step 12:

[0093] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[0094] Step 13:

[0095] The server periodically analyzes the dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, enabling it to provide more appropriate responses in subsequent dialogues.

[0096] Example 1

[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0098] With conventional AI dialogue systems, it can be difficult for users to select a conversation partner who matches their desired characteristics. Furthermore, interactions during a dialogue are not effectively saved and analyzed, making it difficult to improve the system for the next dialogue. This leaves users with a lack of satisfaction and immersion in the system, leading to the need for personalized advice.

[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0100] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved session log and improving response patterns, means for transmitting text data from the user terminal to the server, and means for the server to generate responses based on the extracted keywords, thereby enabling users to communicate in real time with the interlocutor best suited to them.

[0101] "User terminal" means a device (e.g., smartphone, tablet, or PC) used by a user to operate the interface.

[0102] A "server" is a computer system that performs various processes such as text analysis, characteristic extraction, AI personality generation, log storage, and analysis.

[0103] A "database" is a system for storing dialogue logs and training data.

[0104] The "means for accepting user input" is a means for providing an interface for the user to input the characteristics of the interlocutor he / she desires.

[0105] "Morphological analysis" is a method of breaking down input text data into morphemes (the smallest units of language) and analyzing their meaning.

[0106] "Characteristic extraction" is the process of extracting important keywords related to the characteristics of the interlocutor from the analyzed text data.

[0107] An "AI personality" is an artificial intelligence character with specific personality traits and response patterns.

[0108] An "interactive session" refers to a series of interactions between a user and a generated AI personality.

[0109] A "session log" is data that records all interactions during an interactive session.

[0110] "Response pattern" refers to the pattern of responses that an AI personality generates in response to a specific input.

[0111] "Training data" is a dataset used to improve an AI model using historical data such as dialogue logs.

[0112] "Natural language processing technology" is a computer technology for analyzing, understanding, and generating human language.

[0113] A "generative AI model" is an artificial intelligence model that is trained to generate appropriate responses based on user input.

[0114] A "morpheme" is the smallest unit of language, an independent word or its equivalent that has meaning.

[0115] This invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user.

[0116] The system consists of the following components:

[0117] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0118] Server: A computer system that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[0119] Database: A system for storing dialogue logs and training data

[0120] In this system, a user launches the application on a user device such as a smartphone, tablet, or PC. The device displays the question, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, if the user inputs "someone who is quiet and gives logical advice," the device sends this text data to the server.

[0121] The server performs morphological analysis on the received text data to extract important keywords such as "quiet," "logical," and "advice." Natural language processing technology (e.g., NLTK, SpaCy, etc.) is used for morphological analysis. The server then generates an AI personality based on the analyzed data. Specifically, it uses a generative AI model (e.g., GPT-3.5) to select and customize an AI model with response patterns that match specific personality traits.

[0122] The server creates a list of the generated AI personalities and sends it to the device. The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen, and the user selects the desired interlocutor. Once the dialogue session begins, the server and device work together to enable real-time two-way communication. The device sends the questions and consultation details entered by the user to the server, and the server uses the AI ​​personality to generate an appropriate response and send it back to the device. By repeating this process, a natural conversation progresses.

[0123] All messages and exchanges during a conversation session are stored in a database by the server as a conversation log. The server periodically analyzes the saved conversation log and uses it as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses being provided in subsequent conversations.

[0124] Specific examples

[0125] For example, a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics. This input data is sent to the server, and morphological analysis is used to extract the keywords "quiet," "logical," and "advice." Based on this, the server generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other interlocutors, and Erica selects "Ken." When the dialogue session begins and Erica asks about "what's bothering her at work today," the server generates Ken's response and provides appropriate advice. This dialogue log is saved on the server and used for future dialogues.

[0126] Prompt Sentence Examples

[0127] User: Who do you want to talk to today?

[0128] AI: A quiet, logical adviser

[0129] User: Ken (quiet and logical)

[0130] AI: Hi Erica. What can we help you with today?

[0131] User: I'm having trouble with something at work today.

[0132] AI: What specifically are you worried about?

[0133] This system provides users with the most suitable interlocutor and enables real-time communication, thereby providing users with a high level of satisfaction and immersion.

[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0135] Explanation of the system program processing flow and each processing step

[0136] Step 1:

[0137] A user operates a user device (smartphone, tablet, PC) to launch an application. The input is the user's action (tapping the app icon), and the output is the application's home screen displayed on the device. The operation is that the user device displays the application's logo for a few seconds, and then displays the home screen.

[0138] Step 2:

[0139] The terminal displays a question form to the user asking, "Who do you want to talk to today?" The input is a question template in the system, and the output is the question form displayed to the user. Specifically, the terminal displays a form containing a text box and a submit button on the screen.

[0140] Step 3:

[0141] The user inputs the desired characteristics of the interlocutor. For example, the user might input "someone who is quiet and gives logical advice." The input is the user's text data, and the output is the text data sent to the terminal by pressing the send button. The action is for the user to input the characteristics in the text box and click the send button.

[0142] Step 4:

[0143] The terminal sends the input text data to the server. The input is the user's text data, and the output is the text data sent to the server. In operation, the terminal encodes the user data into a packet format and sends it to the server via the Internet.

[0144] Step 5:

[0145] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." The input is the text data sent from the device, and the output is the extracted keywords. In operation, the server analyzes the text data using natural language processing technology (e.g., NLTK, SpaCy, etc.) and breaks it down into morphemes.

[0146] Step 6:

[0147] The server generates an appropriate AI personality based on the extracted keywords. The input is the keywords extracted through analysis, and the output is an AI personality generated by a generative AI model (e.g., GPT-3.5). Specifically, the server uses the trained generative AI model to customize an AI model with response patterns that match specific personality traits.

[0148] Step 7:

[0149] The server creates a list of multiple AI personalities and sends it to the terminal. The input is a list of the generated AI personalities, and the output is the list sent to the user's terminal. In operation, the server packages the information about the AI ​​personalities into a list and sends it to the terminal.

[0150] Step 8:

[0151] The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The input is a list of AI personalities sent from the server, and the output is a list of candidates displayed on the screen. Specifically, the device displays a pop-up window and presents multiple AI personalities as options.

[0152] Step 9:

[0153] The user selects the desired AI personality (e.g., Ken). The input is the user's selection action, and the output is information about the selected AI personality. The user taps the desired interlocutor from the options on the screen.

[0154] Step 10:

[0155] The device notifies the server of the user's selection, and the server initializes an interactive session with the selected AI personality. The input is the user's selection data, and the output is the initial settings for the interactive session. Specifically, the device sends the selection data in packet form to the server, and the server sets up the interactive session.

[0156] Step 11:

[0157] The terminal displays a dialogue window, displaying a first-time greeting and question from the AI ​​interlocutor. The input is the initial response data from the server, and the output is a chat window and greeting message displayed to the user. Specifically, the terminal opens a chat window and displays a message such as "Hello, what would you like to discuss with us today?"

[0158] Step 12:

[0159] The user enters their question or inquiry into the terminal, which then sends it to the server. The input is the user's text data, and the output is the data to be sent to the server. The operation is that the user enters a question in the chat window and clicks the send button.

[0160] Step 13:

[0161] The server generates an appropriate response based on the data sent from the user's device and sends it back to the device. The input is the user's consultation data, and the output is the response generated by the server. Specifically, the server uses a generative AI model to create an appropriate response and sends it to the user's device.

[0162] Step 14:

[0163] The terminal displays the response sent from the server to the user. The input is the response data from the server, and the output is a text message displayed to the user. In operation, the terminal displays the response from the server in a chat window. This process is repeated, achieving real-time two-way communication.

[0164] Step 15:

[0165] The server stores all messages and exchanges during an interactive session as an interaction log in a database. The input is all exchanges during the interaction, and the output is the interaction log stored in the database. Specifically, the server collects all messages in real time and stores them in the database.

[0166] Step 16:

[0167] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the response patterns of the AI ​​personality. The input is the stored dialogue logs, and the output is the improved response patterns. Specifically, the server analyzes the dialogue logs and uses them as a training dataset for the generative AI model. The improved model is reflected in the next session.

[0168] (Application example 1)

[0169] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0170] Today's consumers increasingly desire the support of virtual assistants that can provide specific advice and recommendations, helping them make faster and more accurate product selections. However, many existing virtual assistants are unable to fully address the individual needs and preferences of users and often only provide monotonous responses. This has led to a demand for the development of virtual assistant systems that can provide more personalized assistance to users.

[0171] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0172] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, and means for generating an AI personality as a shopping assistant tailored to the user's individual preferences when the user accesses a virtual store and supporting the user with shopping. This allows the user to proceed with shopping while interacting in real time with a shopping assistant tailored to the user's individual preferences.

[0173] "Means for accepting user input" refers to devices or software that have an interface or function for inputting the characteristics and requests of the desired interlocutor into the system.

[0174] "Means for morphologically analyzing text data entered by a user" refers to natural language processing technology for analyzing text data entered by a user and identifying its components.

[0175] The "means for extracting personality traits from analyzed text data" refers to a function or algorithm for extracting the traits of the interlocutor that the user desires from the data obtained by morphological analysis.

[0176] "Means for generating an AI personality based on extracted personality traits" refers to a mechanism or software that generates an AI personality with specific characteristics based on analyzed information.

[0177] "Means for presenting the generated AI personality to the user" refers to an interface or function that displays options for the generated AI personality to the user and encourages them to make a selection.

[0178] "Means for initiating an interactive session with a user-selected AI personality" refers to the functionality or protocols for initiating two-way communication with a user-selected AI personality.

[0179] The "means for storing interactions during an interactive session" refers to a technique for recording all interactions that occur during an interactive session and storing them for later reference.

[0180] "Means for analyzing stored dialogue logs and improving response patterns" refers to functions and algorithms that analyze stored dialogue logs and use the information obtained from them to improve the AI's response patterns.

[0181] "Means of generating an AI personality as a shopping assistant tailored to individual preferences when a user accesses a virtual store and supporting the user in their shopping" refers to functions and software that automatically generate a shopping assistant tailored to the characteristics and requests of the user when the user uses a virtual store, and support the shopping process.

[0182] This invention is a system that provides immersive communication to users, and in particular, demonstrates its application as a shopping assistant in a virtual store. This system uses specific hardware and software to generate an AI personality based on the characteristics of the person the user desires to interact with, and to realize real-time two-way communication.

[0183] System configuration

[0184] User terminal

[0185] Hardware: Smartphones, tablets, computers, etc.

[0186] Software: Web browser or dedicated application

[0187] server

[0188] Hardware: High-performance server machine

[0189] software:

[0190] Natural Language Processing (NLP) libraries (e.g., spaCy)

[0191] Generative AI models (e.g., GPT-2 by OpenAI)

[0192] Database (e.g. PostgreSQL)

[0193] Web server (e.g. Nginx, Gunicorn)

[0194] Processing flow

[0195] 1. User Input

[0196] The user starts a dedicated application using the user terminal. The application displays a question to the user: "What kind of person do you want to talk to today?" The user then inputs the characteristics of the person they want to talk to. For example, they might input "someone who is knowledgeable about fashion and can tell me about trends."

[0197] 2. Sending and analyzing input data

[0198] The device sends the entered text data to a server, which then receives the data and performs morphological analysis using an NLP library to extract important keywords such as "fashion" and "trend."

[0199] 3. AI personality generation

[0200] The server generates an appropriate AI personality using a generative AI model (e.g., GPT-2) based on the extracted keywords. A prompt such as "You are a shopping assistant with a fashion and trend-oriented personality. Please provide the user with the best advice." is used.

[0201] 4. Starting an interactive session

[0202] The generated AI personalities are presented to the user as multiple options. A dialogue session is initiated with the AI ​​personality selected by the user, and real-time two-way communication takes place. Questions and inquiries entered by the user are sent to the server, which generates an appropriate response and returns it to the user. This process is repeated.

[0203] 5. Saving and analyzing conversation logs

[0204] The server stores logs of all interaction sessions in a database, which are periodically analyzed and used as training data to provide better responses in subsequent interactions.

[0205] Specific examples

[0206] For example, a user named Erika launches a dedicated application and enters "someone who is knowledgeable about fashion and can tell me about trends" as the desired characteristic of her interlocutor. The server analyzes this input data and generates the optimal AI personality. The device presents the generated AI personality to Erika, who selects "Fashion Assistant." When the dialogue session begins and Erika asks, "What are the trends this season?", the server generates a response for the AI ​​assistant and provides appropriate trend information. This dialogue log is saved on the server and used for future dialogues.

[0207] An example prompt is:

[0208] You are a shopping assistant with a fashion and trend-oriented nature. Give your customers the best advice.

[0209] As a result of the above, this invention allows users to proceed with their shopping while interacting in real time with a shopping assistant that is tailored to their individual preferences, resulting in a more satisfying shopping experience.

[0210] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0211] Step 1:

[0212] Input: The user starts the dedicated application using the user terminal.

[0213] What it does: The device displays the application home screen and a question form asking, "Who do you want to talk to today?"

[0214] Output: The user inputs the desired characteristics of the interlocutor.

[0215] Step 2:

[0216] Input: The user inputs, for example, "someone who is knowledgeable about fashion and can tell me about trends" as the characteristics of the person he or she desires to talk to.

[0217] Operation: The terminal sends the entered text data to the server.

[0218] Output: The input data is sent to the server.

[0219] Step 3:

[0220] Input: Text data received by the server

[0221] How it works: The server uses a natural language processing (NLP) library (e.g., spaCy) to perform morphological analysis and extract keywords such as "fashion" and "trend."

[0222] Output: Keywords extracted as analysis results

[0223] Step 4:

[0224] Input: Keywords extracted as analysis results

[0225] How it works: The server uses a generative AI model (e.g., GPT-2) to generate an appropriate AI personality. The prompt used is, "You are a shopping assistant with a fashion and trending personality. Please provide the user with the best advice."

[0226] Output: Generated AI personality

[0227] Step 5:

[0228] Input: Generated AI personality

[0229] How it works: The server creates a list of the generated AI personalities as multiple options and sends it to the device.

[0230] Output: AI personality options displayed on the device

[0231] Step 6:

[0232] Input: AI personality options displayed on the device

[0233] How it works: The user selects the AI ​​personality they want, for example "Fashion Assistant."

[0234] Output: User-selected AI personality

[0235] Step 7:

[0236] Input: User-selected AI personality

[0237] Operation: The device starts an interactive session with the selected AI personality. The server prepares the necessary initial data and initializes the interactive session.

[0238] Output: Interactive session started

[0239] Step 8:

[0240] Input: User's question or inquiry

[0241] How it works: The device sends the question or inquiry entered by the user to the server, which then generates an appropriate response. This response is generated using a generative AI model. For example, if a user asks, "What's trending this season?", the server generates a response and sends it back to the device.

[0242] Output: The AI ​​assistant's response that is displayed to the user

[0243] Step 9:

[0244] Input: All messages and exchanges during an interactive session

[0245] How it works: The server stores all messages and exchanges from an interactive session in a database.

[0246] Output: Saved interaction logs

[0247] Step 10:

[0248] Input: Saved conversation log

[0249] How it works: The server periodically analyzes the stored dialogue logs and uses them as training data to refine the AI ​​personality's response patterns, which will result in more appropriate responses being provided in future dialogues.

[0250] Output: Improved response pattern

[0251] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0252] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses. Below, as an embodiment of the present invention, the processing of the system's program is described in detail in natural language.

[0253] Basic system configuration

[0254] The system consists of the following elements:

[0255] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0256] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[0257] Database: A database that stores interaction logs and training data.

[0258] Emotion Engine: Technology for analyzing and recognizing user emotions

[0259] System processing flow

[0260] 1. The user launches the application

[0261] The user taps the application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0262] 2. Accept user input

[0263] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0264] 3. Sending input data

[0265] The terminal transmits the input text data to the server.

[0266] 4. Morphological analysis and feature extraction

[0267] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0268] 5. AI personality generation

[0269] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0270] 6. User Sentiment Analysis

[0271] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0272] 7. Present interlocutor options

[0273] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0274] 8. Starting an interactive session

[0275] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0276] 9. Real-time two-way communication

[0277] The device sends the user's input questions or consultation details (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[0278] 10. Saving conversation logs

[0279] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0280] 11. Analyzing conversation logs and improving response patterns

[0281] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0282] Specific examples

[0283] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the option of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what's bothering her at work today," the emotion engine analyzes Erica's emotion as "anxiety," and the server generates Ken's response. Ken offers kind words and logical solutions to alleviate Erica's anxiety. This conversation log is saved on the server and used for future conversations.

[0284] As described above, the present invention allows a user to communicate in real time with the person who best suits him or her at any time, and to obtain a satisfying experience that is tailored to his or her emotions.

[0285] The processing flow will be explained below.

[0286] Step 1:

[0287] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0288] Step 2:

[0289] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0290] Step 3:

[0291] The terminal transmits the input text data to the server.

[0292] Step 4:

[0293] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0294] Step 5:

[0295] The server generates an AI personality with the corresponding personality traits based on the extracted keywords, specifically selecting and customizing an AI model with a calm and logical response pattern.

[0296] Step 6:

[0297] The emotion engine installed on the server analyzes emotions from the text entered by the user, determining emotions such as "anxiety," "joy," and "anger."

[0298] Step 7:

[0299] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[0300] Step 8:

[0301] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[0302] Step 9:

[0303] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0304] Step 10:

[0305] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[0306] Step 11:

[0307] The server considers the user's emotions analyzed by the emotion engine and generates a response from the AI ​​personality "Ken." For example, if the user's emotion is recognized as "anxiety," Ken will provide reassuring advice in a calm tone.

[0308] Step 12:

[0309] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[0310] Step 13:

[0311] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[0312] Step 14:

[0313] The server periodically analyzes the saved dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, providing more appropriate responses in subsequent dialogues.

[0314] Example 2

[0315] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0316] Conventional dialogue systems can only provide fixed responses to user input, making it difficult to generate responses that take the user's emotions and circumstances into account. This has led to the issue of not being able to provide a satisfying dialogue experience for users. Furthermore, the learning function required to appropriately utilize dialogue logs and generate more appropriate responses for the next dialogue has also been insufficient.

[0317] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0318] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data input by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, and means for generating responses based on the analyzed emotion information. This makes it possible to generate appropriate responses in real time according to the user's emotions and situation, and to provide more appropriate responses in the next interaction by utilizing the interaction log.

[0319] "User" means any person or entity that interacts with the System.

[0320] "Means for accepting input" refers to an interface that allows a user to input information into the system, and is a system that includes input devices such as a keyboard, touch screen, or voice input.

[0321] "Morphological analysis" is a technique that divides text into words and morphemes and analyzes the meaning and grammatical information of each word.

[0322] The "means for extracting characteristics" is a function for identifying and extracting important keywords and characteristics from morphologically analyzed text data.

[0323] An "artificial intelligence personality" is a virtual conversational agent generated based on extracted characteristics, capable of engaging in natural conversations with users.

[0324] "Means for presentation" refers to a method for visually or audibly displaying the personality of the generated AI to the user, and includes screen display, audio announcement, etc.

[0325] "Interaction session" refers to a series of interactions between a user and a selected AI personality, including multiple message exchanges.

[0326] "Means for storing interactions" refers to a function that records messages and information exchanged between the user and the AI ​​during an interaction session and stores them for later use.

[0327] An "interaction log" is a record of all messages and actions exchanged during an interaction session.

[0328] "Means for improving response patterns" is a function that analyzes saved dialogue logs and learns and improves the AI's responses so that it can respond more appropriately in the next dialogue.

[0329] "Means for analyzing emotions" refers to technology for reading and identifying emotions (e.g., joy, anger, sadness, etc.) from the user's input text or voice.

[0330] The "means for generating a response" is a function that allows the AI ​​personality to create an appropriate response based on the analyzed emotional information and extracted characteristics.

[0331] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses accordingly.

[0332] The basic system configuration is as follows:

[0333] Device: The device through which the user operates the interface (e.g., smartphone, tablet, PC).

[0334] Server: A server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, and sentiment analysis).

[0335] Database: A database that stores dialogue logs and training data.

[0336] Emotion engine: Technology for analyzing and recognizing user emotions.

[0337] The server analyzes the text data entered by the user using a morphological analysis tool called "MeCab," for example, and extracts important keywords. Based on these keywords, the server generates an AI personality using a generative AI model (e.g., GPT-3). Furthermore, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions and adjusts responses based on the analysis results.

[0338] As a concrete example, a user named Erica launches the application and inputs the desired characteristics of a partner: "Someone who is quiet and gives logical advice." The server that receives this input performs morphological analysis and generates an AI personality called "Ken" using GPT-3 based on the extracted keywords. The device then presents "Ken" and other partner options to Erica.

[0339] When Erica selects "Ken" and a dialogue session begins, the server uses the emotion engine to analyze Erica's input and determine her emotions. For example, if Erica asks about her worries at work today, the emotion engine analyzes her emotion as "anxiety." In response, the server generates a response from "Ken" that offers kind words and logical solutions to alleviate Erica's anxiety.

[0340] This dialogue log is stored in a database and analyzed by the server, which uses it as learning data to provide more appropriate responses in subsequent dialogues.

[0341] An example prompt is, "When Erica is feeling anxious about work today, how would Ken, the calm and logical AI personality, respond?"

[0342] In this way, the present invention allows users to communicate with the person who best suits them in real time at any time, providing a satisfying experience tailored to their emotions.

[0343] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0344] Step 1:

[0345] The user launches an application.

[0346] A user taps an application icon on their smartphone to launch the application. The input is the user's action (selecting and tapping the application). The output is the device displaying the application's home screen, which includes the message "Hello! Who would you like to talk to today?"

[0347] Step 2:

[0348] The device displays the input form.

[0349] The device displays a question form to the user asking, "What kind of person do you want to talk to today?" The required input information is the user's action (launching the app) that triggers the display of the form. The output is a text box where the user can enter the characteristics of the person they want to talk to. For example, the user might enter, "Someone who is quiet and gives logical advice."

[0350] Step 3:

[0351] The terminal transmits the text data entered by the user to the server.

[0352] The terminal sends the previously entered text data to the server. The input data is the characteristic information entered by the user. As an output, the sent data arrives at the server and the next processing step is initiated.

[0353] Step 4:

[0354] The server performs morphological analysis on the received text data.

[0355] The server performs morphological analysis of the input text data using an analysis tool (e.g., "MeCab"). The input text data is divided into words and morphemes, and the meaning and grammatical information of each word is analyzed. Important keywords such as "quiet," "logical," and "advice" are extracted as output.

[0356] Step 5:

[0357] The server generates an AI personality based on the data after characteristic extraction.

[0358] The server generates an AI personality using a generative AI model (e.g., GPT-3) based on the extracted keywords. The input is the characteristic keywords extracted through morphological analysis. The output is, for example, "A quiet and logical AI personality named 'Ken.'"

[0359] Step 6:

[0360] The server analyzes the user's emotions.

[0361] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from the user's input text or voice. The input data is the user's text or voice, and the analysis results in emotion labels such as "anger," "joy," or "sadness." In this example, the user's text, "Work wasn't going well today," is analyzed as "anxiety."

[0362] Step 7:

[0363] The server generates multiple AI personalities and sends them to the device.

[0364] The server sends the generated AI personality (e.g., "Ken (quiet and logical)" or "Akira (emotional and encouraging)") to the terminal, which then presents it to the user. The terminal has data for each generated AI personality as input, and displays options to the user as output. For example, the options "Ken" and "Akira" are displayed to the user.

[0365] Step 8:

[0366] The terminal notifies the server of the user's selection, initiating an interactive session.

[0367] The user selects the desired interlocutor, and the terminal notifies the server of the selection. The input is the user's selection action. The output is the initialization of the interaction session, and the terminal displays an initial greeting or question. For example, the message "Hello, I'm Ken! How's your day?" is displayed on the terminal.

[0368] Step 9:

[0369] The user inputs a question or inquiry, and the server generates a response accordingly.

[0370] The user inputs the specific content of their problem (e.g., "What are you worried about at work today?"), and the device sends that data to the server. The input data is the user's problem content and an emotion label. The server receives this and generates an appropriate response based on the emotion analysis results. For example, the output response might be, "That's tough. Can you tell me specifically what happened?"

[0371] Step 10:

[0372] The server logs the interactive session.

[0373] The server stores all messages and exchanges during a conversation session in a database as a conversation log. The input data are messages and emotion labels during the conversation session. The output is the stored log entries that will be used to refine the next response pattern.

[0374] Step 11:

[0375] The server analyzes the dialogue logs and refines response patterns.

[0376] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns. The input data is the stored dialogue logs, and the output is improved response patterns to improve user satisfaction in subsequent dialogue sessions. For example, based on Erica's dialogue history, the system can learn her preferences and tendencies and provide more accurate advice.

[0377] (Application example 2)

[0378] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0379] Conventional dialogue systems have difficulty generating appropriate responses based on the user's emotions, making it impossible to provide dialogue that meets individual needs. Furthermore, when interacting with robots in factories or workplaces, there is a need to provide advice and support that takes into account the worker's emotions. In this situation, it is necessary to develop a system that incorporates emotion analysis.

[0380] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0381] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, means for generating responses according to the analyzed emotions, and means for presenting the responses to the user via a support device. This enables appropriate interaction that takes emotions into consideration.

[0382] "User" refers to the person who operates and interacts with the system.

[0383] "Input" refers to the text data and voice data that the user provides to the system.

[0384] "Means" refers to a method or apparatus for performing a particular function within a system.

[0385] "Morphological analysis" refers to the process of breaking down text data into words and analyzing them.

[0386] "Personality traits" refer to specific personality traits necessary for generating an AI personality.

[0387] "AI personality" refers to a personality model created by artificial intelligence that can converse with users.

[0388] A "session" refers to a series of conversations between a user and an AI personality.

[0389] "Interaction Log" means a record of all messages exchanged during an interaction session.

[0390] "Analysis" refers to the process of finding new information and areas for improvement based on stored data.

[0391] "Emotion" refers to the psychological state that a user exhibits during a conversation.

[0392] "Analysis" refers to the process of examining data in detail to understand its contents.

[0393] "Response" refers to the answer or reaction that the system gives back to the user.

[0394] "Support device" refers to a device that presents the output of the system to the user.

[0395] "Presentation" refers to the act of displaying system-generated information or responses to the user.

[0396] This invention is a system that provides immersive communication to users and can recognize their emotions and adjust responses accordingly. This system generates an AI personality based on the user's needs and engages in real-time dialogue.

[0397] Basic system configuration

[0398] The system consists of the following elements:

[0399] User terminal: The device through which the user operates the interface (e.g., smart glasses, smartphone, computer)

[0400] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[0401] Database: A database that stores interaction logs and training data.

[0402] Emotion Engine: Technology for analyzing and recognizing user emotions

[0403] Hardware and software used

[0404] Hardware: smart glasses (e.g. Microsoft HoloLens), smartphone, PC

[0405] Software: Python, transformers library, sentiment analysis model (EmoRoBERTa), text generation model (GPT-3.5-turbo)

[0406] Processing flow

[0407] 1. The user launches the application

[0408] The user taps the application icon on their smart glasses or smartphone to launch the application, and the device displays the application's home screen.

[0409] 2. Accept user input

[0410] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0411] 3. Sending input data

[0412] The terminal transmits the input text data to the server.

[0413] 4. Morphological analysis and feature extraction

[0414] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0415] 5. AI personality generation

[0416] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0417] 6. User sentiment analysis

[0418] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0419] 7. Present interlocutor options

[0420] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0421] 8. Starting an interactive session

[0422] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0423] 9. Real-time two-way communication

[0424] The device sends the question or consultation content entered by the user (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[0425] 10. Saving conversation logs

[0426] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0427] 11. Analyzing conversation logs and improving response patterns

[0428] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0429] Specific examples

[0430] For example, if a user launches the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics, the server will analyze this input data and generate the optimal AI personality. The device will present the user with several options, and the user will select one. When the dialogue session begins and the user complains that "the machine is not working properly," the emotion engine will interpret this as "anxiety," and the server will generate a response accordingly ("Don't worry, check the instructions in the operating manual and try restarting it as directed") and display it on the smart glasses.

[0431] Prompt Sentence Examples

[0432] Example prompt: "The worker is feeling anxious about the current task. Please consider his / her feelings and provide appropriate advice."

[0433] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0434] Step 1:

[0435] The user launches the application.

[0436] Specific behavior:

[0437] The user turns on the smart glasses or smartphone and taps the application icon.

[0438] input:

[0439] Tap the application icon

[0440] output:

[0441] The application home screen will appear on your device.

[0442] Step 2:

[0443] Accepts user input.

[0444] Specific behavior:

[0445] The device displays a form asking the user, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0446] input:

[0447] Characteristics of desired interlocutor input by user

[0448] output:

[0449] The entered characteristics are recorded as text data.

[0450] Step 3:

[0451] Sending input data.

[0452] Specific behavior:

[0453] The terminal transmits the input text data to the server.

[0454] input:

[0455] Text data entered by the user

[0456] output:

[0457] The text data is sent to the server.

[0458] Step 4:

[0459] Morphological analysis and feature extraction.

[0460] Specific behavior:

[0461] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0462] input:

[0463] Text data sent to the server

[0464] output:

[0465] Important keywords are extracted.

[0466] Step 5:

[0467] AI personality generation.

[0468] Specific behavior:

[0469] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it uses a generative AI model to select and customize an AI personality with a calm and logical response pattern.

[0470] input:

[0471] Extracted keywords

[0472] output:

[0473] Generated AI personality

[0474] Step 6:

[0475] User sentiment analysis.

[0476] Specific behavior:

[0477] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0478] input:

[0479] User input text or voice

[0480] output:

[0481] Emotion Labels

[0482] Step 7:

[0483] Presenting interlocutor options.

[0484] Specific behavior:

[0485] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0486] input:

[0487] List of generated AI personalities

[0488] output:

[0489] The list of contacts presented to the user

[0490] Step 8:

[0491] Start an interactive session.

[0492] Specific behavior:

[0493] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0494] input:

[0495] Notification of the user's choice of AI personality

[0496] output:

[0497] Displaying a dialogue window, initial greetings and questions

[0498] Step 9:

[0499] Real-time two-way communication.

[0500] Specific behavior:

[0501] The device sends the questions and consultation details entered by the user to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[0502] input:

[0503] User questions, consultation details, and emotion labels

[0504] output:

[0505] Emotion-aware responses

[0506] Step 10:

[0507] Saving conversation logs.

[0508] Specific behavior:

[0509] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0510] input:

[0511] Interactive session messages and exchanges

[0512] output:

[0513] Interaction logs stored in a database

[0514] Step 11:

[0515] Analyzing dialogue logs and refining response patterns.

[0516] Specific behavior:

[0517] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0518] input:

[0519] Saved conversation logs

[0520] output:

[0521] Improved response patterns

[0522] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0523] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0524] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0525] [Second embodiment]

[0526] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0527] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0528] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0529] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0530] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0531] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0532] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0533] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0534] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0535] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0536] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0537] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0538] The present invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user. Below, as an embodiment of the present invention, the processing of the system's program is explained in detail in natural language.

[0539] Basic system configuration

[0540] The system consists of the following elements:

[0541] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0542] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[0543] Database: A database that stores interaction logs and training data.

[0544] System processing flow

[0545] 1. The user launches the application

[0546] The user taps the application icon on their smartphone to launch it, and the device displays the application's home screen.

[0547] 2. Accept user input

[0548] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, the user might input, "Someone who is quiet and gives logical advice."

[0549] 3. Sending input data

[0550] The terminal transmits the input text data to the server.

[0551] 4. Morphological analysis and feature extraction

[0552] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." This analysis is performed using natural language processing technology.

[0553] 5. AI personality generation

[0554] The server generates an appropriate AI personality based on the extracted keywords, for example, selecting and customizing an AI model with a response pattern that matches a specific personality trait (e.g., quiet and logical).

[0555] 6. Present interlocutor options

[0556] The server creates a list of multiple AI personalities and sends it to the device. The device then displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The user then selects the desired interlocutor.

[0557] 7. Starting an interactive session

[0558] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0559] 8. Real-time two-way communication

[0560] The device sends the questions or inquiries entered by the user to the server, which then uses its AI personality to generate an appropriate response and sends it back to the device, which then displays it to the user. This process is repeated, enabling real-time two-way communication.

[0561] 9. Saving conversation logs

[0562] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0563] 10. Analyzing conversation logs and improving response patterns

[0564] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0565] Specific examples

[0566] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what is bothering her at work today," the server generates Ken's response and provides appropriate advice. This conversation log is saved on the server and used for future conversations.

[0567] As described above, the present invention allows users to communicate in real time with the interlocutor that best suits them at any time, providing an immersive and satisfying experience.

[0568] The processing flow will be explained below.

[0569] Step 1:

[0570] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0571] Step 2:

[0572] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0573] Step 3:

[0574] The terminal transmits the input text data to the server.

[0575] Step 4:

[0576] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0577] Step 5:

[0578] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0579] Step 6:

[0580] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[0581] Step 7:

[0582] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[0583] Step 8:

[0584] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0585] Step 9:

[0586] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[0587] Step 10:

[0588] The server generates a response from the AI ​​personality "Ken" and sends it back to the device. For example, the server generates logical advice as text.

[0589] Step 11:

[0590] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[0591] Step 12:

[0592] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[0593] Step 13:

[0594] The server periodically analyzes the dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, enabling it to provide more appropriate responses in subsequent dialogues.

[0595] Example 1

[0596] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0597] With conventional AI dialogue systems, it can be difficult for users to select a conversation partner who matches their desired characteristics. Furthermore, interactions during a dialogue are not effectively saved and analyzed, making it difficult to improve the system for the next dialogue. This leaves users with a lack of satisfaction and immersion in the system, leading to the need for personalized advice.

[0598] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0599] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved session log and improving response patterns, means for transmitting text data from the user terminal to the server, and means for the server to generate responses based on the extracted keywords, thereby enabling users to communicate in real time with the interlocutor best suited to them.

[0600] "User terminal" means a device (e.g., smartphone, tablet, or PC) used by a user to operate the interface.

[0601] A "server" is a computer system that performs various processes such as text analysis, characteristic extraction, AI personality generation, log storage, and analysis.

[0602] A "database" is a system for storing dialogue logs and training data.

[0603] The "means for accepting user input" is a means for providing an interface for the user to input the characteristics of the interlocutor he / she desires.

[0604] "Morphological analysis" is a method of breaking down input text data into morphemes (the smallest units of language) and analyzing their meaning.

[0605] "Characteristic extraction" is the process of extracting important keywords related to the characteristics of the interlocutor from the analyzed text data.

[0606] An "AI personality" is an artificial intelligence character with specific personality traits and response patterns.

[0607] An "interactive session" refers to a series of interactions between a user and a generated AI personality.

[0608] A "session log" is data that records all interactions during an interactive session.

[0609] "Response pattern" refers to the pattern of responses that an AI personality generates in response to a specific input.

[0610] "Training data" is a dataset used to improve an AI model using historical data such as dialogue logs.

[0611] "Natural language processing technology" is a computer technology for analyzing, understanding, and generating human language.

[0612] A "generative AI model" is an artificial intelligence model that is trained to generate appropriate responses based on user input.

[0613] A "morpheme" is the smallest unit of language, an independent word or its equivalent that has meaning.

[0614] This invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user.

[0615] The system consists of the following components:

[0616] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0617] Server: A computer system that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[0618] Database: A system for storing dialogue logs and training data

[0619] In this system, a user launches the application on a user device such as a smartphone, tablet, or PC. The device displays the question, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, if the user inputs "someone who is quiet and gives logical advice," the device sends this text data to the server.

[0620] The server performs morphological analysis on the received text data to extract important keywords such as "quiet," "logical," and "advice." Natural language processing technology (e.g., NLTK, SpaCy, etc.) is used for morphological analysis. The server then generates an AI personality based on the analyzed data. Specifically, it uses a generative AI model (e.g., GPT-3.5) to select and customize an AI model with response patterns that match specific personality traits.

[0621] The server creates a list of the generated AI personalities and sends it to the device. The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen, and the user selects the desired interlocutor. Once the dialogue session begins, the server and device work together to enable real-time two-way communication. The device sends the questions and consultation details entered by the user to the server, and the server uses the AI ​​personality to generate an appropriate response and send it back to the device. By repeating this process, a natural conversation progresses.

[0622] All messages and exchanges during a conversation session are stored in a database by the server as a conversation log. The server periodically analyzes the saved conversation log and uses it as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses being provided in subsequent conversations.

[0623] Specific examples

[0624] For example, a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics. This input data is sent to the server, and morphological analysis is used to extract the keywords "quiet," "logical," and "advice." Based on this, the server generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other interlocutors, and Erica selects "Ken." When the dialogue session begins and Erica asks about "what's bothering her at work today," the server generates Ken's response and provides appropriate advice. This dialogue log is saved on the server and used for future dialogues.

[0625] Prompt Sentence Examples

[0626] User: Who do you want to talk to today?

[0627] AI: A quiet, logical adviser

[0628] User: Ken (quiet and logical)

[0629] AI: Hi Erica. What can we help you with today?

[0630] User: I'm having trouble with something at work today.

[0631] AI: What specifically are you worried about?

[0632] This system provides users with the most suitable interlocutor and enables real-time communication, thereby providing users with a high level of satisfaction and immersion.

[0633] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0634] Explanation of the system program processing flow and each processing step

[0635] Step 1:

[0636] A user operates a user device (smartphone, tablet, PC) to launch an application. The input is the user's action (tapping the app icon), and the output is the application's home screen displayed on the device. The operation is that the user device displays the application's logo for a few seconds, and then displays the home screen.

[0637] Step 2:

[0638] The terminal displays a question form to the user asking, "Who do you want to talk to today?" The input is a question template in the system, and the output is the question form displayed to the user. Specifically, the terminal displays a form containing a text box and a submit button on the screen.

[0639] Step 3:

[0640] The user inputs the desired characteristics of the interlocutor. For example, the user might input "someone who is quiet and gives logical advice." The input is the user's text data, and the output is the text data sent to the terminal by pressing the send button. The action is for the user to input the characteristics in the text box and click the send button.

[0641] Step 4:

[0642] The terminal sends the input text data to the server. The input is the user's text data, and the output is the text data sent to the server. In operation, the terminal encodes the user data into a packet format and sends it to the server via the Internet.

[0643] Step 5:

[0644] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." The input is the text data sent from the device, and the output is the extracted keywords. In operation, the server analyzes the text data using natural language processing technology (e.g., NLTK, SpaCy, etc.) and breaks it down into morphemes.

[0645] Step 6:

[0646] The server generates an appropriate AI personality based on the extracted keywords. The input is the keywords extracted through analysis, and the output is an AI personality generated by a generative AI model (e.g., GPT-3.5). Specifically, the server uses the trained generative AI model to customize an AI model with response patterns that match specific personality traits.

[0647] Step 7:

[0648] The server creates a list of multiple AI personalities and sends it to the terminal. The input is a list of the generated AI personalities, and the output is the list sent to the user's terminal. In operation, the server packages the information about the AI ​​personalities into a list and sends it to the terminal.

[0649] Step 8:

[0650] The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The input is a list of AI personalities sent from the server, and the output is a list of candidates displayed on the screen. Specifically, the device displays a pop-up window and presents multiple AI personalities as options.

[0651] Step 9:

[0652] The user selects the desired AI personality (e.g., Ken). The input is the user's selection action, and the output is information about the selected AI personality. The user taps the desired interlocutor from the options on the screen.

[0653] Step 10:

[0654] The device notifies the server of the user's selection, and the server initializes an interactive session with the selected AI personality. The input is the user's selection data, and the output is the initial settings for the interactive session. Specifically, the device sends the selection data in packet form to the server, and the server sets up the interactive session.

[0655] Step 11:

[0656] The terminal displays a dialogue window, displaying a first-time greeting and question from the AI ​​interlocutor. The input is the initial response data from the server, and the output is a chat window and greeting message displayed to the user. Specifically, the terminal opens a chat window and displays a message such as "Hello, what would you like to discuss with us today?"

[0657] Step 12:

[0658] The user enters their question or inquiry into the terminal, which then sends it to the server. The input is the user's text data, and the output is the data to be sent to the server. The operation is that the user enters a question in the chat window and clicks the send button.

[0659] Step 13:

[0660] The server generates an appropriate response based on the data sent from the user's device and sends it back to the device. The input is the user's consultation data, and the output is the response generated by the server. Specifically, the server uses a generative AI model to create an appropriate response and sends it to the user's device.

[0661] Step 14:

[0662] The terminal displays the response sent from the server to the user. The input is the response data from the server, and the output is a text message displayed to the user. In operation, the terminal displays the response from the server in a chat window. This process is repeated, achieving real-time two-way communication.

[0663] Step 15:

[0664] The server stores all messages and exchanges during an interactive session as an interaction log in a database. The input is all exchanges during the interaction, and the output is the interaction log stored in the database. Specifically, the server collects all messages in real time and stores them in the database.

[0665] Step 16:

[0666] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the response patterns of the AI ​​personality. The input is the stored dialogue logs, and the output is the improved response patterns. Specifically, the server analyzes the dialogue logs and uses them as a training dataset for the generative AI model. The improved model is reflected in the next session.

[0667] (Application example 1)

[0668] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0669] Today's consumers increasingly desire the support of virtual assistants that can provide specific advice and recommendations, helping them make faster and more accurate product selections. However, many existing virtual assistants are unable to fully address the individual needs and preferences of users and often only provide monotonous responses. This has led to a demand for the development of virtual assistant systems that can provide more personalized assistance to users.

[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0671] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, and means for generating an AI personality as a shopping assistant tailored to the user's individual preferences when the user accesses a virtual store and supporting the user with shopping. This allows the user to proceed with shopping while interacting in real time with a shopping assistant tailored to the user's individual preferences.

[0672] "Means for accepting user input" refers to devices or software that have an interface or function for inputting the characteristics and requests of the desired interlocutor into the system.

[0673] "Means for morphologically analyzing text data entered by a user" refers to natural language processing technology for analyzing text data entered by a user and identifying its components.

[0674] The "means for extracting personality traits from analyzed text data" refers to a function or algorithm for extracting the traits of the interlocutor that the user desires from the data obtained by morphological analysis.

[0675] "Means for generating an AI personality based on extracted personality traits" refers to a mechanism or software that generates an AI personality with specific characteristics based on analyzed information.

[0676] "Means for presenting the generated AI personality to the user" refers to an interface or function that displays options for the generated AI personality to the user and encourages them to make a selection.

[0677] "Means for initiating an interactive session with a user-selected AI personality" refers to the functionality or protocols for initiating two-way communication with a user-selected AI personality.

[0678] The "means for storing interactions during an interactive session" refers to a technique for recording all interactions that occur during an interactive session and storing them for later reference.

[0679] "Means for analyzing stored dialogue logs and improving response patterns" refers to functions and algorithms that analyze stored dialogue logs and use the information obtained from them to improve the AI's response patterns.

[0680] "Means of generating an AI personality as a shopping assistant tailored to individual preferences when a user accesses a virtual store and supporting the user in their shopping" refers to functions and software that automatically generate a shopping assistant tailored to the characteristics and requests of the user when the user uses a virtual store, and support the shopping process.

[0681] This invention is a system that provides immersive communication to users, and in particular, demonstrates its application as a shopping assistant in a virtual store. This system uses specific hardware and software to generate an AI personality based on the characteristics of the person the user desires to interact with, and to realize real-time two-way communication.

[0682] System configuration

[0683] User terminal

[0684] Hardware: Smartphones, tablets, computers, etc.

[0685] Software: Web browser or dedicated application

[0686] server

[0687] Hardware: High-performance server machine

[0688] software:

[0689] Natural Language Processing (NLP) libraries (e.g., spaCy)

[0690] Generative AI models (e.g., GPT-2 by OpenAI)

[0691] Database (e.g. PostgreSQL)

[0692] Web server (e.g. Nginx, Gunicorn)

[0693] Processing flow

[0694] 1. User Input

[0695] The user starts a dedicated application using the user terminal. The application displays a question to the user: "What kind of person do you want to talk to today?" The user then inputs the characteristics of the person they want to talk to. For example, they might input "someone who is knowledgeable about fashion and can tell me about trends."

[0696] 2. Sending and analyzing input data

[0697] The device sends the entered text data to a server, which then receives the data and performs morphological analysis using an NLP library to extract important keywords such as "fashion" and "trend."

[0698] 3. AI personality generation

[0699] The server generates an appropriate AI personality using a generative AI model (e.g., GPT-2) based on the extracted keywords. A prompt such as "You are a shopping assistant with a fashion and trend-oriented personality. Please provide the user with the best advice." is used.

[0700] 4. Starting an interactive session

[0701] The generated AI personalities are presented to the user as multiple options. A dialogue session is initiated with the AI ​​personality selected by the user, and real-time two-way communication takes place. Questions and inquiries entered by the user are sent to the server, which generates an appropriate response and returns it to the user. This process is repeated.

[0702] 5. Saving and analyzing conversation logs

[0703] The server stores logs of all interaction sessions in a database, which are periodically analyzed and used as training data to provide better responses in subsequent interactions.

[0704] Specific examples

[0705] For example, a user named Erika launches a dedicated application and enters "someone who is knowledgeable about fashion and can tell me about trends" as the desired characteristic of her interlocutor. The server analyzes this input data and generates the optimal AI personality. The device presents the generated AI personality to Erika, who selects "Fashion Assistant." When the dialogue session begins and Erika asks, "What are the trends this season?", the server generates a response for the AI ​​assistant and provides appropriate trend information. This dialogue log is saved on the server and used for future dialogues.

[0706] An example prompt is:

[0707] You are a shopping assistant with a fashion and trend-oriented nature. Give your customers the best advice.

[0708] As a result of the above, this invention allows users to proceed with their shopping while interacting in real time with a shopping assistant that is tailored to their individual preferences, resulting in a more satisfying shopping experience.

[0709] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0710] Step 1:

[0711] Input: The user starts the dedicated application using the user terminal.

[0712] What it does: The device displays the application home screen and a question form asking, "Who do you want to talk to today?"

[0713] Output: The user inputs the desired characteristics of the interlocutor.

[0714] Step 2:

[0715] Input: The user inputs, for example, "someone who is knowledgeable about fashion and can tell me about trends" as the characteristics of the person he or she desires to talk to.

[0716] Operation: The terminal sends the entered text data to the server.

[0717] Output: The input data is sent to the server.

[0718] Step 3:

[0719] Input: Text data received by the server

[0720] How it works: The server uses a natural language processing (NLP) library (e.g., spaCy) to perform morphological analysis and extract keywords such as "fashion" and "trend."

[0721] Output: Keywords extracted as analysis results

[0722] Step 4:

[0723] Input: Keywords extracted as analysis results

[0724] How it works: The server uses a generative AI model (e.g., GPT-2) to generate an appropriate AI personality. The prompt used is, "You are a shopping assistant with a fashion and trending personality. Please provide the user with the best advice."

[0725] Output: Generated AI personality

[0726] Step 5:

[0727] Input: Generated AI personality

[0728] How it works: The server creates a list of the generated AI personalities as multiple options and sends it to the device.

[0729] Output: AI personality options displayed on the device

[0730] Step 6:

[0731] Input: AI personality options displayed on the device

[0732] How it works: The user selects the AI ​​personality they want, for example "Fashion Assistant."

[0733] Output: User-selected AI personality

[0734] Step 7:

[0735] Input: User-selected AI personality

[0736] Operation: The device starts an interactive session with the selected AI personality. The server prepares the necessary initial data and initializes the interactive session.

[0737] Output: Interactive session started

[0738] Step 8:

[0739] Input: User's question or inquiry

[0740] How it works: The device sends the question or inquiry entered by the user to the server, which then generates an appropriate response. This response is generated using a generative AI model. For example, if a user asks, "What's trending this season?", the server generates a response and sends it back to the device.

[0741] Output: The AI ​​assistant's response that is displayed to the user

[0742] Step 9:

[0743] Input: All messages and exchanges during an interactive session

[0744] How it works: The server stores all messages and exchanges from an interactive session in a database.

[0745] Output: Saved interaction logs

[0746] Step 10:

[0747] Input: Saved conversation log

[0748] How it works: The server periodically analyzes the stored dialogue logs and uses them as training data to refine the AI ​​personality's response patterns, which will result in more appropriate responses being provided in future dialogues.

[0749] Output: Improved response pattern

[0750] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0751] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses. Below, as an embodiment of the present invention, the processing of the system's program is described in detail in natural language.

[0752] Basic system configuration

[0753] The system consists of the following elements:

[0754] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[0755] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[0756] Database: A database that stores interaction logs and training data.

[0757] Emotion Engine: Technology for analyzing and recognizing user emotions

[0758] System processing flow

[0759] 1. The user launches the application

[0760] The user taps the application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0761] 2. Accept user input

[0762] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0763] 3. Sending input data

[0764] The terminal transmits the input text data to the server.

[0765] 4. Morphological analysis and feature extraction

[0766] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0767] 5. AI personality generation

[0768] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0769] 6. User Sentiment Analysis

[0770] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0771] 7. Present interlocutor options

[0772] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0773] 8. Starting an interactive session

[0774] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0775] 9. Real-time two-way communication

[0776] The device sends the user's input questions or consultation details (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[0777] 10. Saving conversation logs

[0778] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0779] 11. Analyzing conversation logs and improving response patterns

[0780] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0781] Specific examples

[0782] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the option of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what's bothering her at work today," the emotion engine analyzes Erica's emotion as "anxiety," and the server generates Ken's response. Ken offers kind words and logical solutions to alleviate Erica's anxiety. This conversation log is saved on the server and used for future conversations.

[0783] As described above, the present invention allows a user to communicate in real time with the person who best suits him or her at any time, and to obtain a satisfying experience that is tailored to his or her emotions.

[0784] The processing flow will be explained below.

[0785] Step 1:

[0786] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[0787] Step 2:

[0788] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0789] Step 3:

[0790] The terminal transmits the input text data to the server.

[0791] Step 4:

[0792] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0793] Step 5:

[0794] The server generates an AI personality with the corresponding personality traits based on the extracted keywords, specifically selecting and customizing an AI model with a calm and logical response pattern.

[0795] Step 6:

[0796] The emotion engine installed on the server analyzes emotions from the text entered by the user, determining emotions such as "anxiety," "joy," and "anger."

[0797] Step 7:

[0798] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[0799] Step 8:

[0800] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[0801] Step 9:

[0802] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0803] Step 10:

[0804] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[0805] Step 11:

[0806] The server considers the user's emotions analyzed by the emotion engine and generates a response from the AI ​​personality "Ken." For example, if the user's emotion is recognized as "anxiety," Ken will provide reassuring advice in a calm tone.

[0807] Step 12:

[0808] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[0809] Step 13:

[0810] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[0811] Step 14:

[0812] The server periodically analyzes the saved dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, providing more appropriate responses in subsequent dialogues.

[0813] Example 2

[0814] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0815] Conventional dialogue systems can only provide fixed responses to user input, making it difficult to generate responses that take the user's emotions and circumstances into account. This has led to the issue of not being able to provide a satisfying dialogue experience for users. Furthermore, the learning function required to appropriately utilize dialogue logs and generate more appropriate responses for the next dialogue has also been insufficient.

[0816] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0817] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data input by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, and means for generating responses based on the analyzed emotion information. This makes it possible to generate appropriate responses in real time according to the user's emotions and situation, and to provide more appropriate responses in the next interaction by utilizing the interaction log.

[0818] "User" means any person or entity that interacts with the System.

[0819] "Means for accepting input" refers to an interface that allows a user to input information into the system, and is a system that includes input devices such as a keyboard, touch screen, or voice input.

[0820] "Morphological analysis" is a technique that divides text into words and morphemes and analyzes the meaning and grammatical information of each word.

[0821] The "means for extracting characteristics" is a function for identifying and extracting important keywords and characteristics from morphologically analyzed text data.

[0822] An "artificial intelligence personality" is a virtual conversational agent generated based on extracted characteristics, capable of engaging in natural conversations with users.

[0823] "Means for presentation" refers to a method for visually or audibly displaying the personality of the generated AI to the user, and includes screen display, audio announcement, etc.

[0824] "Interaction session" refers to a series of interactions between a user and a selected AI personality, including multiple message exchanges.

[0825] "Means for storing interactions" refers to a function that records messages and information exchanged between the user and the AI ​​during an interaction session and stores them for later use.

[0826] An "interaction log" is a record of all messages and actions exchanged during an interaction session.

[0827] "Means for improving response patterns" is a function that analyzes saved dialogue logs and learns and improves the AI's responses so that it can respond more appropriately in the next dialogue.

[0828] "Means for analyzing emotions" refers to technology for reading and identifying emotions (e.g., joy, anger, sadness, etc.) from the user's input text or voice.

[0829] The "means for generating a response" is a function that allows the AI ​​personality to create an appropriate response based on the analyzed emotional information and extracted characteristics.

[0830] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses accordingly.

[0831] The basic system configuration is as follows:

[0832] Device: The device through which the user operates the interface (e.g., smartphone, tablet, PC).

[0833] Server: A server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, and sentiment analysis).

[0834] Database: A database that stores dialogue logs and training data.

[0835] Emotion engine: Technology for analyzing and recognizing user emotions.

[0836] The server analyzes the text data entered by the user using a morphological analysis tool called "MeCab," for example, and extracts important keywords. Based on these keywords, the server generates an AI personality using a generative AI model (e.g., GPT-3). Furthermore, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions and adjusts responses based on the analysis results.

[0837] As a concrete example, a user named Erica launches the application and inputs the desired characteristics of a partner: "Someone who is quiet and gives logical advice." The server that receives this input performs morphological analysis and generates an AI personality called "Ken" using GPT-3 based on the extracted keywords. The device then presents "Ken" and other partner options to Erica.

[0838] When Erica selects "Ken" and a dialogue session begins, the server uses the emotion engine to analyze Erica's input and determine her emotions. For example, if Erica asks about her worries at work today, the emotion engine analyzes her emotion as "anxiety." In response, the server generates a response from "Ken" that offers kind words and logical solutions to alleviate Erica's anxiety.

[0839] This dialogue log is stored in a database and analyzed by the server, which uses it as learning data to provide more appropriate responses in subsequent dialogues.

[0840] An example prompt is, "When Erica is feeling anxious about work today, how would Ken, the calm and logical AI personality, respond?"

[0841] In this way, the present invention allows users to communicate with the person who best suits them in real time at any time, providing a satisfying experience tailored to their emotions.

[0842] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0843] Step 1:

[0844] The user launches an application.

[0845] A user taps an application icon on their smartphone to launch the application. The input is the user's action (selecting and tapping the application). The output is the device displaying the application's home screen, which includes the message "Hello! Who would you like to talk to today?"

[0846] Step 2:

[0847] The device displays the input form.

[0848] The device displays a question form to the user asking, "What kind of person do you want to talk to today?" The required input information is the user's action (launching the app) that triggers the display of the form. The output is a text box where the user can enter the characteristics of the person they want to talk to. For example, the user might enter, "Someone who is quiet and gives logical advice."

[0849] Step 3:

[0850] The terminal transmits the text data entered by the user to the server.

[0851] The terminal sends the previously entered text data to the server. The input data is the characteristic information entered by the user. As an output, the sent data arrives at the server and the next processing step is initiated.

[0852] Step 4:

[0853] The server performs morphological analysis on the received text data.

[0854] The server performs morphological analysis of the input text data using an analysis tool (e.g., "MeCab"). The input text data is divided into words and morphemes, and the meaning and grammatical information of each word is analyzed. Important keywords such as "quiet," "logical," and "advice" are extracted as output.

[0855] Step 5:

[0856] The server generates an AI personality based on the data after characteristic extraction.

[0857] The server generates an AI personality using a generative AI model (e.g., GPT-3) based on the extracted keywords. The input is the characteristic keywords extracted through morphological analysis. The output is, for example, "A quiet and logical AI personality named 'Ken.'"

[0858] Step 6:

[0859] The server analyzes the user's emotions.

[0860] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from the user's input text or voice. The input data is the user's text or voice, and the analysis results in emotion labels such as "anger," "joy," or "sadness." In this example, the user's text, "Work wasn't going well today," is analyzed as "anxiety."

[0861] Step 7:

[0862] The server generates multiple AI personalities and sends them to the device.

[0863] The server sends the generated AI personality (e.g., "Ken (quiet and logical)" or "Akira (emotional and encouraging)") to the terminal, which then presents it to the user. The terminal has data for each generated AI personality as input, and displays options to the user as output. For example, the options "Ken" and "Akira" are displayed to the user.

[0864] Step 8:

[0865] The terminal notifies the server of the user's selection, initiating an interactive session.

[0866] The user selects the desired interlocutor, and the terminal notifies the server of the selection. The input is the user's selection action. The output is the initialization of the interaction session, and the terminal displays an initial greeting or question. For example, the message "Hello, I'm Ken! How's your day?" is displayed on the terminal.

[0867] Step 9:

[0868] The user inputs a question or inquiry, and the server generates a response accordingly.

[0869] The user inputs the specific content of their problem (e.g., "What are you worried about at work today?"), and the device sends that data to the server. The input data is the user's problem content and an emotion label. The server receives this and generates an appropriate response based on the emotion analysis results. For example, the output response might be, "That's tough. Can you tell me specifically what happened?"

[0870] Step 10:

[0871] The server logs the interactive session.

[0872] The server stores all messages and exchanges during a conversation session in a database as a conversation log. The input data are messages and emotion labels during the conversation session. The output is the stored log entries that will be used to refine the next response pattern.

[0873] Step 11:

[0874] The server analyzes the dialogue logs and refines response patterns.

[0875] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns. The input data is the stored dialogue logs, and the output is improved response patterns to improve user satisfaction in subsequent dialogue sessions. For example, based on Erica's dialogue history, the system can learn her preferences and tendencies and provide more accurate advice.

[0876] (Application example 2)

[0877] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0878] Conventional dialogue systems have difficulty generating appropriate responses based on the user's emotions, making it impossible to provide dialogue that meets individual needs. Furthermore, when interacting with robots in factories or workplaces, there is a need to provide advice and support that takes into account the worker's emotions. In this situation, it is necessary to develop a system that incorporates emotion analysis.

[0879] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0880] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, means for generating responses according to the analyzed emotions, and means for presenting the responses to the user via a support device. This enables appropriate interaction that takes emotions into consideration.

[0881] "User" refers to the person who operates and interacts with the system.

[0882] "Input" refers to the text data and voice data that the user provides to the system.

[0883] "Means" refers to a method or apparatus for performing a particular function within a system.

[0884] "Morphological analysis" refers to the process of breaking down text data into words and analyzing them.

[0885] "Personality traits" refer to specific personality traits necessary for generating an AI personality.

[0886] "AI personality" refers to a personality model created by artificial intelligence that can converse with users.

[0887] A "session" refers to a series of conversations between a user and an AI personality.

[0888] "Interaction Log" means a record of all messages exchanged during an interaction session.

[0889] "Analysis" refers to the process of finding new information and areas for improvement based on stored data.

[0890] "Emotion" refers to the psychological state that a user exhibits during a conversation.

[0891] "Analysis" refers to the process of examining data in detail to understand its contents.

[0892] "Response" refers to the answer or reaction that the system gives back to the user.

[0893] "Support device" refers to a device that presents the output of the system to the user.

[0894] "Presentation" refers to the act of displaying system-generated information or responses to the user.

[0895] This invention is a system that provides immersive communication to users and can recognize their emotions and adjust responses accordingly. This system generates an AI personality based on the user's needs and engages in real-time dialogue.

[0896] Basic system configuration

[0897] The system consists of the following elements:

[0898] User terminal: The device through which the user operates the interface (e.g., smart glasses, smartphone, computer)

[0899] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[0900] Database: A database that stores interaction logs and training data.

[0901] Emotion Engine: Technology for analyzing and recognizing user emotions

[0902] Hardware and software used

[0903] Hardware: smart glasses (e.g. Microsoft HoloLens), smartphone, PC

[0904] Software: Python, transformers library, sentiment analysis model (EmoRoBERTa), text generation model (GPT-3.5-turbo)

[0905] Processing flow

[0906] 1. The user launches the application

[0907] The user taps the application icon on their smart glasses or smartphone to launch the application, and the device displays the application's home screen.

[0908] 2. Accept user input

[0909] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0910] 3. Sending input data

[0911] The terminal transmits the input text data to the server.

[0912] 4. Morphological analysis and feature extraction

[0913] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0914] 5. AI personality generation

[0915] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[0916] 6. User sentiment analysis

[0917] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0918] 7. Present interlocutor options

[0919] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0920] 8. Starting an interactive session

[0921] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0922] 9. Real-time two-way communication

[0923] The device sends the question or consultation content entered by the user (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[0924] 10. Saving conversation logs

[0925] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[0926] 11. Analyzing conversation logs and improving response patterns

[0927] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[0928] Specific examples

[0929] For example, if a user launches the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics, the server will analyze this input data and generate the optimal AI personality. The device will present the user with several options, and the user will select one. When the dialogue session begins and the user complains that "the machine is not working properly," the emotion engine will interpret this as "anxiety," and the server will generate a response accordingly ("Don't worry, check the instructions in the operating manual and try restarting it as directed") and display it on the smart glasses.

[0930] Prompt Sentence Examples

[0931] Example prompt: "The worker is feeling anxious about the current task. Please consider his / her feelings and provide appropriate advice."

[0932] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0933] Step 1:

[0934] The user launches the application.

[0935] Specific behavior:

[0936] The user turns on the smart glasses or smartphone and taps the application icon.

[0937] input:

[0938] Tap the application icon

[0939] output:

[0940] The application home screen will appear on your device.

[0941] Step 2:

[0942] Accepts user input.

[0943] Specific behavior:

[0944] The device displays a form asking the user, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[0945] input:

[0946] Characteristics of desired interlocutor input by user

[0947] output:

[0948] The entered characteristics are recorded as text data.

[0949] Step 3:

[0950] Sending input data.

[0951] Specific behavior:

[0952] The terminal transmits the input text data to the server.

[0953] input:

[0954] Text data entered by the user

[0955] output:

[0956] The text data is sent to the server.

[0957] Step 4:

[0958] Morphological analysis and feature extraction.

[0959] Specific behavior:

[0960] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[0961] input:

[0962] Text data sent to the server

[0963] output:

[0964] Important keywords are extracted.

[0965] Step 5:

[0966] AI personality generation.

[0967] Specific behavior:

[0968] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it uses a generative AI model to select and customize an AI personality with a calm and logical response pattern.

[0969] input:

[0970] Extracted keywords

[0971] output:

[0972] Generated AI personality

[0973] Step 6:

[0974] User sentiment analysis.

[0975] Specific behavior:

[0976] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[0977] input:

[0978] User input text or voice

[0979] output:

[0980] Emotion Labels

[0981] Step 7:

[0982] Presenting interlocutor options.

[0983] Specific behavior:

[0984] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[0985] input:

[0986] List of generated AI personalities

[0987] output:

[0988] The list of contacts presented to the user

[0989] Step 8:

[0990] Start an interactive session.

[0991] Specific behavior:

[0992] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[0993] input:

[0994] Notification of the user's choice of AI personality

[0995] output:

[0996] Displaying a dialogue window, initial greetings and questions

[0997] Step 9:

[0998] Real-time two-way communication.

[0999] Specific behavior:

[1000] The device sends the questions and consultation details entered by the user to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1001] input:

[1002] User questions, consultation details, and emotion labels

[1003] output:

[1004] Emotion-aware responses

[1005] Step 10:

[1006] Saving conversation logs.

[1007] Specific behavior:

[1008] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1009] input:

[1010] Interactive session messages and exchanges

[1011] output:

[1012] Interaction logs stored in a database

[1013] Step 11:

[1014] Analyzing dialogue logs and refining response patterns.

[1015] Specific behavior:

[1016] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1017] input:

[1018] Saved conversation logs

[1019] output:

[1020] Improved response patterns

[1021] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1022] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1023] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1024] [Third embodiment]

[1025] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1026] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1028] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1029] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1030] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1032] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1033] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1035] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1036] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1037] The present invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user. Below, as an embodiment of the present invention, the processing of the system's program is explained in detail in natural language.

[1038] Basic system configuration

[1039] The system consists of the following elements:

[1040] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1041] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[1042] Database: A database that stores interaction logs and training data.

[1043] System processing flow

[1044] 1. The user launches the application

[1045] The user taps the application icon on their smartphone to launch it, and the device displays the application's home screen.

[1046] 2. Accept user input

[1047] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, the user might input, "Someone who is quiet and gives logical advice."

[1048] 3. Sending input data

[1049] The terminal transmits the input text data to the server.

[1050] 4. Morphological analysis and feature extraction

[1051] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." This analysis is performed using natural language processing technology.

[1052] 5. AI personality generation

[1053] The server generates an appropriate AI personality based on the extracted keywords, for example, selecting and customizing an AI model with a response pattern that matches a specific personality trait (e.g., quiet and logical).

[1054] 6. Present interlocutor options

[1055] The server creates a list of multiple AI personalities and sends it to the device. The device then displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The user then selects the desired interlocutor.

[1056] 7. Starting an interactive session

[1057] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1058] 8. Real-time two-way communication

[1059] The device sends the questions or inquiries entered by the user to the server, which then uses its AI personality to generate an appropriate response and sends it back to the device, which then displays it to the user. This process is repeated, enabling real-time two-way communication.

[1060] 9. Saving conversation logs

[1061] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1062] 10. Analyzing conversation logs and improving response patterns

[1063] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1064] Specific examples

[1065] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what is bothering her at work today," the server generates Ken's response and provides appropriate advice. This conversation log is saved on the server and used for future conversations.

[1066] As described above, the present invention allows users to communicate in real time with the interlocutor that best suits them at any time, providing an immersive and satisfying experience.

[1067] The processing flow will be explained below.

[1068] Step 1:

[1069] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1070] Step 2:

[1071] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1072] Step 3:

[1073] The terminal transmits the input text data to the server.

[1074] Step 4:

[1075] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1076] Step 5:

[1077] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1078] Step 6:

[1079] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[1080] Step 7:

[1081] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[1082] Step 8:

[1083] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1084] Step 9:

[1085] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[1086] Step 10:

[1087] The server generates a response from the AI ​​personality "Ken" and sends it back to the device. For example, the server generates logical advice as text.

[1088] Step 11:

[1089] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[1090] Step 12:

[1091] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[1092] Step 13:

[1093] The server periodically analyzes the dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, enabling it to provide more appropriate responses in subsequent dialogues.

[1094] Example 1

[1095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1096] With conventional AI dialogue systems, it can be difficult for users to select a conversation partner who matches their desired characteristics. Furthermore, interactions during a dialogue are not effectively saved and analyzed, making it difficult to improve the system for the next dialogue. This leaves users with a lack of satisfaction and immersion in the system, leading to the need for personalized advice.

[1097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1098] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved session log and improving response patterns, means for transmitting text data from the user terminal to the server, and means for the server to generate responses based on the extracted keywords, thereby enabling users to communicate in real time with the interlocutor best suited to them.

[1099] "User terminal" means a device (e.g., smartphone, tablet, or PC) used by a user to operate the interface.

[1100] A "server" is a computer system that performs various processes such as text analysis, characteristic extraction, AI personality generation, log storage, and analysis.

[1101] A "database" is a system for storing dialogue logs and training data.

[1102] The "means for accepting user input" is a means for providing an interface for the user to input the characteristics of the interlocutor he / she desires.

[1103] "Morphological analysis" is a method of breaking down input text data into morphemes (the smallest units of language) and analyzing their meaning.

[1104] "Characteristic extraction" is the process of extracting important keywords related to the characteristics of the interlocutor from the analyzed text data.

[1105] An "AI personality" is an artificial intelligence character with specific personality traits and response patterns.

[1106] An "interactive session" refers to a series of interactions between a user and a generated AI personality.

[1107] A "session log" is data that records all interactions during an interactive session.

[1108] "Response pattern" refers to the pattern of responses that an AI personality generates in response to a specific input.

[1109] "Training data" is a dataset used to improve an AI model using historical data such as dialogue logs.

[1110] "Natural language processing technology" is a computer technology for analyzing, understanding, and generating human language.

[1111] A "generative AI model" is an artificial intelligence model that is trained to generate appropriate responses based on user input.

[1112] A "morpheme" is the smallest unit of language, an independent word or its equivalent that has meaning.

[1113] This invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user.

[1114] The system consists of the following components:

[1115] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1116] Server: A computer system that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[1117] Database: A system for storing dialogue logs and training data

[1118] In this system, a user launches the application on a user device such as a smartphone, tablet, or PC. The device displays the question, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, if the user inputs "someone who is quiet and gives logical advice," the device sends this text data to the server.

[1119] The server performs morphological analysis on the received text data to extract important keywords such as "quiet," "logical," and "advice." Natural language processing technology (e.g., NLTK, SpaCy, etc.) is used for morphological analysis. The server then generates an AI personality based on the analyzed data. Specifically, it uses a generative AI model (e.g., GPT-3.5) to select and customize an AI model with response patterns that match specific personality traits.

[1120] The server creates a list of the generated AI personalities and sends it to the device. The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen, and the user selects the desired interlocutor. Once the dialogue session begins, the server and device work together to enable real-time two-way communication. The device sends the questions and consultation details entered by the user to the server, and the server uses the AI ​​personality to generate an appropriate response and send it back to the device. By repeating this process, a natural conversation progresses.

[1121] All messages and exchanges during a conversation session are stored in a database by the server as a conversation log. The server periodically analyzes the saved conversation log and uses it as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses being provided in subsequent conversations.

[1122] Specific examples

[1123] For example, a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics. This input data is sent to the server, and morphological analysis is used to extract the keywords "quiet," "logical," and "advice." Based on this, the server generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other interlocutors, and Erica selects "Ken." When the dialogue session begins and Erica asks about "what's bothering her at work today," the server generates Ken's response and provides appropriate advice. This dialogue log is saved on the server and used for future dialogues.

[1124] Prompt Sentence Examples

[1125] User: Who do you want to talk to today?

[1126] AI: A quiet, logical adviser

[1127] User: Ken (quiet and logical)

[1128] AI: Hi Erica. What can we help you with today?

[1129] User: I'm having trouble with something at work today.

[1130] AI: What specifically are you worried about?

[1131] This system provides users with the most suitable interlocutor and enables real-time communication, thereby providing users with a high level of satisfaction and immersion.

[1132] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1133] Explanation of the system program processing flow and each processing step

[1134] Step 1:

[1135] A user operates a user device (smartphone, tablet, PC) to launch an application. The input is the user's action (tapping the app icon), and the output is the application's home screen displayed on the device. The operation is that the user device displays the application's logo for a few seconds, and then displays the home screen.

[1136] Step 2:

[1137] The terminal displays a question form to the user asking, "Who do you want to talk to today?" The input is a question template in the system, and the output is the question form displayed to the user. Specifically, the terminal displays a form containing a text box and a submit button on the screen.

[1138] Step 3:

[1139] The user inputs the desired characteristics of the interlocutor. For example, the user might input "someone who is quiet and gives logical advice." The input is the user's text data, and the output is the text data sent to the terminal by pressing the send button. The action is for the user to input the characteristics in the text box and click the send button.

[1140] Step 4:

[1141] The terminal sends the input text data to the server. The input is the user's text data, and the output is the text data sent to the server. In operation, the terminal encodes the user data into a packet format and sends it to the server via the Internet.

[1142] Step 5:

[1143] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." The input is the text data sent from the device, and the output is the extracted keywords. In operation, the server analyzes the text data using natural language processing technology (e.g., NLTK, SpaCy, etc.) and breaks it down into morphemes.

[1144] Step 6:

[1145] The server generates an appropriate AI personality based on the extracted keywords. The input is the keywords extracted through analysis, and the output is an AI personality generated by a generative AI model (e.g., GPT-3.5). Specifically, the server uses the trained generative AI model to customize an AI model with response patterns that match specific personality traits.

[1146] Step 7:

[1147] The server creates a list of multiple AI personalities and sends it to the terminal. The input is a list of the generated AI personalities, and the output is the list sent to the user's terminal. In operation, the server packages the information about the AI ​​personalities into a list and sends it to the terminal.

[1148] Step 8:

[1149] The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The input is a list of AI personalities sent from the server, and the output is a list of candidates displayed on the screen. Specifically, the device displays a pop-up window and presents multiple AI personalities as options.

[1150] Step 9:

[1151] The user selects the desired AI personality (e.g., Ken). The input is the user's selection action, and the output is information about the selected AI personality. The user taps the desired interlocutor from the options on the screen.

[1152] Step 10:

[1153] The device notifies the server of the user's selection, and the server initializes an interactive session with the selected AI personality. The input is the user's selection data, and the output is the initial settings for the interactive session. Specifically, the device sends the selection data in packet form to the server, and the server sets up the interactive session.

[1154] Step 11:

[1155] The terminal displays a dialogue window, displaying a first-time greeting and question from the AI ​​interlocutor. The input is the initial response data from the server, and the output is a chat window and greeting message displayed to the user. Specifically, the terminal opens a chat window and displays a message such as "Hello, what would you like to discuss with us today?"

[1156] Step 12:

[1157] The user enters their question or inquiry into the terminal, which then sends it to the server. The input is the user's text data, and the output is the data to be sent to the server. The operation is that the user enters a question in the chat window and clicks the send button.

[1158] Step 13:

[1159] The server generates an appropriate response based on the data sent from the user's device and sends it back to the device. The input is the user's consultation data, and the output is the response generated by the server. Specifically, the server uses a generative AI model to create an appropriate response and sends it to the user's device.

[1160] Step 14:

[1161] The terminal displays the response sent from the server to the user. The input is the response data from the server, and the output is a text message displayed to the user. In operation, the terminal displays the response from the server in a chat window. This process is repeated, achieving real-time two-way communication.

[1162] Step 15:

[1163] The server stores all messages and exchanges during an interactive session as an interaction log in a database. The input is all exchanges during the interaction, and the output is the interaction log stored in the database. Specifically, the server collects all messages in real time and stores them in the database.

[1164] Step 16:

[1165] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the response patterns of the AI ​​personality. The input is the stored dialogue logs, and the output is the improved response patterns. Specifically, the server analyzes the dialogue logs and uses them as a training dataset for the generative AI model. The improved model is reflected in the next session.

[1166] (Application example 1)

[1167] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1168] Today's consumers increasingly desire the support of virtual assistants that can provide specific advice and recommendations, helping them make faster and more accurate product selections. However, many existing virtual assistants are unable to fully address the individual needs and preferences of users and often only provide monotonous responses. This has led to a demand for the development of virtual assistant systems that can provide more personalized assistance to users.

[1169] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1170] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, and means for generating an AI personality as a shopping assistant tailored to the user's individual preferences when the user accesses a virtual store and supporting the user with shopping. This allows the user to proceed with shopping while interacting in real time with a shopping assistant tailored to the user's individual preferences.

[1171] "Means for accepting user input" refers to devices or software that have an interface or function for inputting the characteristics and requests of the desired interlocutor into the system.

[1172] "Means for morphologically analyzing text data entered by a user" refers to natural language processing technology for analyzing text data entered by a user and identifying its components.

[1173] The "means for extracting personality traits from analyzed text data" refers to a function or algorithm for extracting the traits of the interlocutor that the user desires from the data obtained by morphological analysis.

[1174] "Means for generating an AI personality based on extracted personality traits" refers to a mechanism or software that generates an AI personality with specific characteristics based on analyzed information.

[1175] "Means for presenting the generated AI personality to the user" refers to an interface or function that displays options for the generated AI personality to the user and encourages them to make a selection.

[1176] "Means for initiating an interactive session with a user-selected AI personality" refers to the functionality or protocols for initiating two-way communication with a user-selected AI personality.

[1177] The "means for storing interactions during an interactive session" refers to a technique for recording all interactions that occur during an interactive session and storing them for later reference.

[1178] "Means for analyzing stored dialogue logs and improving response patterns" refers to functions and algorithms that analyze stored dialogue logs and use the information obtained from them to improve the AI's response patterns.

[1179] "Means of generating an AI personality as a shopping assistant tailored to individual preferences when a user accesses a virtual store and supporting the user in their shopping" refers to functions and software that automatically generate a shopping assistant tailored to the characteristics and requests of the user when the user uses a virtual store, and support the shopping process.

[1180] This invention is a system that provides immersive communication to users, and in particular, demonstrates its application as a shopping assistant in a virtual store. This system uses specific hardware and software to generate an AI personality based on the characteristics of the person the user desires to interact with, and to realize real-time two-way communication.

[1181] System configuration

[1182] User terminal

[1183] Hardware: Smartphones, tablets, computers, etc.

[1184] Software: Web browser or dedicated application

[1185] server

[1186] Hardware: High-performance server machine

[1187] software:

[1188] Natural Language Processing (NLP) libraries (e.g., spaCy)

[1189] Generative AI models (e.g., GPT-2 by OpenAI)

[1190] Database (e.g. PostgreSQL)

[1191] Web server (e.g. Nginx, Gunicorn)

[1192] Processing flow

[1193] 1. User Input

[1194] The user starts a dedicated application using the user terminal. The application displays a question to the user: "What kind of person do you want to talk to today?" The user then inputs the characteristics of the person they want to talk to. For example, they might input "someone who is knowledgeable about fashion and can tell me about trends."

[1195] 2. Sending and analyzing input data

[1196] The device sends the entered text data to a server, which then receives the data and performs morphological analysis using an NLP library to extract important keywords such as "fashion" and "trend."

[1197] 3. AI personality generation

[1198] The server generates an appropriate AI personality using a generative AI model (e.g., GPT-2) based on the extracted keywords. A prompt such as "You are a shopping assistant with a fashion and trend-oriented personality. Please provide the user with the best advice." is used.

[1199] 4. Starting an interactive session

[1200] The generated AI personalities are presented to the user as multiple options. A dialogue session is initiated with the AI ​​personality selected by the user, and real-time two-way communication takes place. Questions and inquiries entered by the user are sent to the server, which generates an appropriate response and returns it to the user. This process is repeated.

[1201] 5. Saving and analyzing conversation logs

[1202] The server stores logs of all interaction sessions in a database, which are periodically analyzed and used as training data to provide better responses in subsequent interactions.

[1203] Specific examples

[1204] For example, a user named Erika launches a dedicated application and enters "someone who is knowledgeable about fashion and can tell me about trends" as the desired characteristic of her interlocutor. The server analyzes this input data and generates the optimal AI personality. The device presents the generated AI personality to Erika, who selects "Fashion Assistant." When the dialogue session begins and Erika asks, "What are the trends this season?", the server generates a response for the AI ​​assistant and provides appropriate trend information. This dialogue log is saved on the server and used for future dialogues.

[1205] An example prompt is:

[1206] You are a shopping assistant with a fashion and trend-oriented nature. Give your customers the best advice.

[1207] As a result of the above, this invention allows users to proceed with their shopping while interacting in real time with a shopping assistant that is tailored to their individual preferences, resulting in a more satisfying shopping experience.

[1208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1209] Step 1:

[1210] Input: The user starts the dedicated application using the user terminal.

[1211] What it does: The device displays the application home screen and a question form asking, "Who do you want to talk to today?"

[1212] Output: The user inputs the desired characteristics of the interlocutor.

[1213] Step 2:

[1214] Input: The user inputs, for example, "someone who is knowledgeable about fashion and can tell me about trends" as the characteristics of the person he or she desires to talk to.

[1215] Operation: The terminal sends the entered text data to the server.

[1216] Output: The input data is sent to the server.

[1217] Step 3:

[1218] Input: Text data received by the server

[1219] How it works: The server uses a natural language processing (NLP) library (e.g., spaCy) to perform morphological analysis and extract keywords such as "fashion" and "trend."

[1220] Output: Keywords extracted as analysis results

[1221] Step 4:

[1222] Input: Keywords extracted as analysis results

[1223] How it works: The server uses a generative AI model (e.g., GPT-2) to generate an appropriate AI personality. The prompt used is, "You are a shopping assistant with a fashion and trending personality. Please provide the user with the best advice."

[1224] Output: Generated AI personality

[1225] Step 5:

[1226] Input: Generated AI personality

[1227] How it works: The server creates a list of the generated AI personalities as multiple options and sends it to the device.

[1228] Output: AI personality options displayed on the device

[1229] Step 6:

[1230] Input: AI personality options displayed on the device

[1231] How it works: The user selects the AI ​​personality they want, for example "Fashion Assistant."

[1232] Output: User-selected AI personality

[1233] Step 7:

[1234] Input: User-selected AI personality

[1235] Operation: The device starts an interactive session with the selected AI personality. The server prepares the necessary initial data and initializes the interactive session.

[1236] Output: Interactive session started

[1237] Step 8:

[1238] Input: User's question or inquiry

[1239] How it works: The device sends the question or inquiry entered by the user to the server, which then generates an appropriate response. This response is generated using a generative AI model. For example, if a user asks, "What's trending this season?", the server generates a response and sends it back to the device.

[1240] Output: The AI ​​assistant's response that is displayed to the user

[1241] Step 9:

[1242] Input: All messages and exchanges during an interactive session

[1243] How it works: The server stores all messages and exchanges from an interactive session in a database.

[1244] Output: Saved interaction logs

[1245] Step 10:

[1246] Input: Saved conversation log

[1247] How it works: The server periodically analyzes the stored dialogue logs and uses them as training data to refine the AI ​​personality's response patterns, which will result in more appropriate responses being provided in future dialogues.

[1248] Output: Improved response pattern

[1249] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1250] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses. Below, as an embodiment of the present invention, the processing of the system's program is described in detail in natural language.

[1251] Basic system configuration

[1252] The system consists of the following elements:

[1253] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1254] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[1255] Database: A database that stores interaction logs and training data.

[1256] Emotion Engine: Technology for analyzing and recognizing user emotions

[1257] System processing flow

[1258] 1. The user launches the application

[1259] The user taps the application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1260] 2. Accept user input

[1261] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1262] 3. Sending input data

[1263] The terminal transmits the input text data to the server.

[1264] 4. Morphological analysis and feature extraction

[1265] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1266] 5. AI personality generation

[1267] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1268] 6. User Sentiment Analysis

[1269] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1270] 7. Present interlocutor options

[1271] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1272] 8. Starting an interactive session

[1273] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1274] 9. Real-time two-way communication

[1275] The device sends the user's input questions or consultation details (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1276] 10. Saving conversation logs

[1277] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1278] 11. Analyzing conversation logs and improving response patterns

[1279] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1280] Specific examples

[1281] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the option of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what's bothering her at work today," the emotion engine analyzes Erica's emotion as "anxiety," and the server generates Ken's response. Ken offers kind words and logical solutions to alleviate Erica's anxiety. This conversation log is saved on the server and used for future conversations.

[1282] As described above, the present invention allows a user to communicate in real time with the person who best suits him or her at any time, and to obtain a satisfying experience that is tailored to his or her emotions.

[1283] The processing flow will be explained below.

[1284] Step 1:

[1285] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1286] Step 2:

[1287] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1288] Step 3:

[1289] The terminal transmits the input text data to the server.

[1290] Step 4:

[1291] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1292] Step 5:

[1293] The server generates an AI personality with the corresponding personality traits based on the extracted keywords, specifically selecting and customizing an AI model with a calm and logical response pattern.

[1294] Step 6:

[1295] The emotion engine installed on the server analyzes emotions from the text entered by the user, determining emotions such as "anxiety," "joy," and "anger."

[1296] Step 7:

[1297] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[1298] Step 8:

[1299] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[1300] Step 9:

[1301] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1302] Step 10:

[1303] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[1304] Step 11:

[1305] The server considers the user's emotions analyzed by the emotion engine and generates a response from the AI ​​personality "Ken." For example, if the user's emotion is recognized as "anxiety," Ken will provide reassuring advice in a calm tone.

[1306] Step 12:

[1307] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[1308] Step 13:

[1309] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[1310] Step 14:

[1311] The server periodically analyzes the saved dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, providing more appropriate responses in subsequent dialogues.

[1312] Example 2

[1313] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1314] Conventional dialogue systems can only provide fixed responses to user input, making it difficult to generate responses that take the user's emotions and circumstances into account. This has led to the issue of not being able to provide a satisfying dialogue experience for users. Furthermore, the learning function required to appropriately utilize dialogue logs and generate more appropriate responses for the next dialogue has also been insufficient.

[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1316] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data input by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, and means for generating responses based on the analyzed emotion information. This makes it possible to generate appropriate responses in real time according to the user's emotions and situation, and to provide more appropriate responses in the next interaction by utilizing the interaction log.

[1317] "User" means any person or entity that interacts with the System.

[1318] "Means for accepting input" refers to an interface that allows a user to input information into the system, and is a system that includes input devices such as a keyboard, touch screen, or voice input.

[1319] "Morphological analysis" is a technique that divides text into words and morphemes and analyzes the meaning and grammatical information of each word.

[1320] The "means for extracting characteristics" is a function for identifying and extracting important keywords and characteristics from morphologically analyzed text data.

[1321] An "artificial intelligence personality" is a virtual conversational agent generated based on extracted characteristics, capable of engaging in natural conversations with users.

[1322] "Means for presentation" refers to a method for visually or audibly displaying the personality of the generated AI to the user, and includes screen display, audio announcement, etc.

[1323] "Interaction session" refers to a series of interactions between a user and a selected AI personality, including multiple message exchanges.

[1324] "Means for storing interactions" refers to a function that records messages and information exchanged between the user and the AI ​​during an interaction session and stores them for later use.

[1325] An "interaction log" is a record of all messages and actions exchanged during an interaction session.

[1326] "Means for improving response patterns" is a function that analyzes saved dialogue logs and learns and improves the AI's responses so that it can respond more appropriately in the next dialogue.

[1327] "Means for analyzing emotions" refers to technology for reading and identifying emotions (e.g., joy, anger, sadness, etc.) from the user's input text or voice.

[1328] The "means for generating a response" is a function that allows the AI ​​personality to create an appropriate response based on the analyzed emotional information and extracted characteristics.

[1329] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses accordingly.

[1330] The basic system configuration is as follows:

[1331] Device: The device through which the user operates the interface (e.g., smartphone, tablet, PC).

[1332] Server: A server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, and sentiment analysis).

[1333] Database: A database that stores dialogue logs and training data.

[1334] Emotion engine: Technology for analyzing and recognizing user emotions.

[1335] The server analyzes the text data entered by the user using a morphological analysis tool called "MeCab," for example, and extracts important keywords. Based on these keywords, the server generates an AI personality using a generative AI model (e.g., GPT-3). Furthermore, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions and adjusts responses based on the analysis results.

[1336] As a concrete example, a user named Erica launches the application and inputs the desired characteristics of a partner: "Someone who is quiet and gives logical advice." The server that receives this input performs morphological analysis and generates an AI personality called "Ken" using GPT-3 based on the extracted keywords. The device then presents "Ken" and other partner options to Erica.

[1337] When Erica selects "Ken" and a dialogue session begins, the server uses the emotion engine to analyze Erica's input and determine her emotions. For example, if Erica asks about her worries at work today, the emotion engine analyzes her emotion as "anxiety." In response, the server generates a response from "Ken" that offers kind words and logical solutions to alleviate Erica's anxiety.

[1338] This dialogue log is stored in a database and analyzed by the server, which uses it as learning data to provide more appropriate responses in subsequent dialogues.

[1339] An example prompt is, "When Erica is feeling anxious about work today, how would Ken, the calm and logical AI personality, respond?"

[1340] In this way, the present invention allows users to communicate with the person who best suits them in real time at any time, providing a satisfying experience tailored to their emotions.

[1341] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1342] Step 1:

[1343] The user launches an application.

[1344] A user taps an application icon on their smartphone to launch the application. The input is the user's action (selecting and tapping the application). The output is the device displaying the application's home screen, which includes the message "Hello! Who would you like to talk to today?"

[1345] Step 2:

[1346] The device displays the input form.

[1347] The device displays a question form to the user asking, "What kind of person do you want to talk to today?" The required input information is the user's action (launching the app) that triggers the display of the form. The output is a text box where the user can enter the characteristics of the person they want to talk to. For example, the user might enter, "Someone who is quiet and gives logical advice."

[1348] Step 3:

[1349] The terminal transmits the text data entered by the user to the server.

[1350] The terminal sends the previously entered text data to the server. The input data is the characteristic information entered by the user. As an output, the sent data arrives at the server and the next processing step is initiated.

[1351] Step 4:

[1352] The server performs morphological analysis on the received text data.

[1353] The server performs morphological analysis of the input text data using an analysis tool (e.g., "MeCab"). The input text data is divided into words and morphemes, and the meaning and grammatical information of each word is analyzed. Important keywords such as "quiet," "logical," and "advice" are extracted as output.

[1354] Step 5:

[1355] The server generates an AI personality based on the data after characteristic extraction.

[1356] The server generates an AI personality using a generative AI model (e.g., GPT-3) based on the extracted keywords. The input is the characteristic keywords extracted through morphological analysis. The output is, for example, "A quiet and logical AI personality named 'Ken.'"

[1357] Step 6:

[1358] The server analyzes the user's emotions.

[1359] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from the user's input text or voice. The input data is the user's text or voice, and the analysis results in emotion labels such as "anger," "joy," or "sadness." In this example, the user's text, "Work wasn't going well today," is analyzed as "anxiety."

[1360] Step 7:

[1361] The server generates multiple AI personalities and sends them to the device.

[1362] The server sends the generated AI personality (e.g., "Ken (quiet and logical)" or "Akira (emotional and encouraging)") to the terminal, which then presents it to the user. The terminal has data for each generated AI personality as input, and displays options to the user as output. For example, the options "Ken" and "Akira" are displayed to the user.

[1363] Step 8:

[1364] The terminal notifies the server of the user's selection, initiating an interactive session.

[1365] The user selects the desired interlocutor, and the terminal notifies the server of the selection. The input is the user's selection action. The output is the initialization of the interaction session, and the terminal displays an initial greeting or question. For example, the message "Hello, I'm Ken! How's your day?" is displayed on the terminal.

[1366] Step 9:

[1367] The user inputs a question or inquiry, and the server generates a response accordingly.

[1368] The user inputs the specific content of their problem (e.g., "What are you worried about at work today?"), and the device sends that data to the server. The input data is the user's problem content and an emotion label. The server receives this and generates an appropriate response based on the emotion analysis results. For example, the output response might be, "That's tough. Can you tell me specifically what happened?"

[1369] Step 10:

[1370] The server logs the interactive session.

[1371] The server stores all messages and exchanges during a conversation session in a database as a conversation log. The input data are messages and emotion labels during the conversation session. The output is the stored log entries that will be used to refine the next response pattern.

[1372] Step 11:

[1373] The server analyzes the dialogue logs and refines response patterns.

[1374] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns. The input data is the stored dialogue logs, and the output is improved response patterns to improve user satisfaction in subsequent dialogue sessions. For example, based on Erica's dialogue history, the system can learn her preferences and tendencies and provide more accurate advice.

[1375] (Application example 2)

[1376] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1377] Conventional dialogue systems have difficulty generating appropriate responses based on the user's emotions, making it impossible to provide dialogue that meets individual needs. Furthermore, when interacting with robots in factories or workplaces, there is a need to provide advice and support that takes into account the worker's emotions. In this situation, it is necessary to develop a system that incorporates emotion analysis.

[1378] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1379] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, means for generating responses according to the analyzed emotions, and means for presenting the responses to the user via a support device. This enables appropriate interaction that takes emotions into consideration.

[1380] "User" refers to the person who operates and interacts with the system.

[1381] "Input" refers to the text data and voice data that the user provides to the system.

[1382] "Means" refers to a method or apparatus for performing a particular function within a system.

[1383] "Morphological analysis" refers to the process of breaking down text data into words and analyzing them.

[1384] "Personality traits" refer to specific personality traits necessary for generating an AI personality.

[1385] "AI personality" refers to a personality model created by artificial intelligence that can converse with users.

[1386] A "session" refers to a series of conversations between a user and an AI personality.

[1387] "Interaction Log" means a record of all messages exchanged during an interaction session.

[1388] "Analysis" refers to the process of finding new information and areas for improvement based on stored data.

[1389] "Emotion" refers to the psychological state that a user exhibits during a conversation.

[1390] "Analysis" refers to the process of examining data in detail to understand its contents.

[1391] "Response" refers to the answer or reaction that the system gives back to the user.

[1392] "Support device" refers to a device that presents the output of the system to the user.

[1393] "Presentation" refers to the act of displaying system-generated information or responses to the user.

[1394] This invention is a system that provides immersive communication to users and can recognize their emotions and adjust responses accordingly. This system generates an AI personality based on the user's needs and engages in real-time dialogue.

[1395] Basic system configuration

[1396] The system consists of the following elements:

[1397] User terminal: The device through which the user operates the interface (e.g., smart glasses, smartphone, computer)

[1398] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[1399] Database: A database that stores interaction logs and training data.

[1400] Emotion Engine: Technology for analyzing and recognizing user emotions

[1401] Hardware and software used

[1402] Hardware: smart glasses (e.g. Microsoft HoloLens), smartphone, PC

[1403] Software: Python, transformers library, sentiment analysis model (EmoRoBERTa), text generation model (GPT-3.5-turbo)

[1404] Processing flow

[1405] 1. The user launches the application

[1406] The user taps the application icon on their smart glasses or smartphone to launch the application, and the device displays the application's home screen.

[1407] 2. Accept user input

[1408] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1409] 3. Sending input data

[1410] The terminal transmits the input text data to the server.

[1411] 4. Morphological analysis and feature extraction

[1412] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1413] 5. AI personality generation

[1414] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1415] 6. User sentiment analysis

[1416] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1417] 7. Present interlocutor options

[1418] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1419] 8. Starting an interactive session

[1420] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1421] 9. Real-time two-way communication

[1422] The device sends the question or consultation content entered by the user (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1423] 10. Saving conversation logs

[1424] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1425] 11. Analyzing conversation logs and improving response patterns

[1426] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1427] Specific examples

[1428] For example, if a user launches the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics, the server will analyze this input data and generate the optimal AI personality. The device will present the user with several options, and the user will select one. When the dialogue session begins and the user complains that "the machine is not working properly," the emotion engine will interpret this as "anxiety," and the server will generate a response accordingly ("Don't worry, check the instructions in the operating manual and try restarting it as directed") and display it on the smart glasses.

[1429] Prompt Sentence Examples

[1430] Example prompt: "The worker is feeling anxious about the current task. Please consider his / her feelings and provide appropriate advice."

[1431] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1432] Step 1:

[1433] The user launches the application.

[1434] Specific behavior:

[1435] The user turns on the smart glasses or smartphone and taps the application icon.

[1436] input:

[1437] Tap the application icon

[1438] output:

[1439] The application home screen will appear on your device.

[1440] Step 2:

[1441] Accepts user input.

[1442] Specific behavior:

[1443] The device displays a form asking the user, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1444] input:

[1445] Characteristics of desired interlocutor input by user

[1446] output:

[1447] The entered characteristics are recorded as text data.

[1448] Step 3:

[1449] Sending input data.

[1450] Specific behavior:

[1451] The terminal transmits the input text data to the server.

[1452] input:

[1453] Text data entered by the user

[1454] output:

[1455] The text data is sent to the server.

[1456] Step 4:

[1457] Morphological analysis and feature extraction.

[1458] Specific behavior:

[1459] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1460] input:

[1461] Text data sent to the server

[1462] output:

[1463] Important keywords are extracted.

[1464] Step 5:

[1465] AI personality generation.

[1466] Specific behavior:

[1467] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it uses a generative AI model to select and customize an AI personality with a calm and logical response pattern.

[1468] input:

[1469] Extracted keywords

[1470] output:

[1471] Generated AI personality

[1472] Step 6:

[1473] User sentiment analysis.

[1474] Specific behavior:

[1475] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1476] input:

[1477] User input text or voice

[1478] output:

[1479] Emotion Labels

[1480] Step 7:

[1481] Presenting interlocutor options.

[1482] Specific behavior:

[1483] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1484] input:

[1485] List of generated AI personalities

[1486] output:

[1487] The list of contacts presented to the user

[1488] Step 8:

[1489] Start an interactive session.

[1490] Specific behavior:

[1491] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1492] input:

[1493] Notification of the user's choice of AI personality

[1494] output:

[1495] Displaying a dialogue window, initial greetings and questions

[1496] Step 9:

[1497] Real-time two-way communication.

[1498] Specific behavior:

[1499] The device sends the questions and consultation details entered by the user to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1500] input:

[1501] User questions, consultation details, and emotion labels

[1502] output:

[1503] Emotion-aware responses

[1504] Step 10:

[1505] Saving conversation logs.

[1506] Specific behavior:

[1507] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1508] input:

[1509] Interactive session messages and exchanges

[1510] output:

[1511] Interaction logs stored in a database

[1512] Step 11:

[1513] Analyzing dialogue logs and refining response patterns.

[1514] Specific behavior:

[1515] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1516] input:

[1517] Saved conversation logs

[1518] output:

[1519] Improved response patterns

[1520] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1521] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1522] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1523] [Fourth embodiment]

[1524] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1525] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1526] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1527] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1528] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1529] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1530] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1531] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1532] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1533] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1534] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1535] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1536] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1537] The present invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user. Below, as an embodiment of the present invention, the processing of the system's program is explained in detail in natural language.

[1538] Basic system configuration

[1539] The system consists of the following elements:

[1540] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1541] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[1542] Database: A database that stores interaction logs and training data.

[1543] System processing flow

[1544] 1. The user launches the application

[1545] The user taps the application icon on their smartphone to launch it, and the device displays the application's home screen.

[1546] 2. Accept user input

[1547] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, the user might input, "Someone who is quiet and gives logical advice."

[1548] 3. Sending input data

[1549] The terminal transmits the input text data to the server.

[1550] 4. Morphological analysis and feature extraction

[1551] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." This analysis is performed using natural language processing technology.

[1552] 5. AI personality generation

[1553] The server generates an appropriate AI personality based on the extracted keywords, for example, selecting and customizing an AI model with a response pattern that matches a specific personality trait (e.g., quiet and logical).

[1554] 6. Present interlocutor options

[1555] The server creates a list of multiple AI personalities and sends it to the device. The device then displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The user then selects the desired interlocutor.

[1556] 7. Starting an interactive session

[1557] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1558] 8. Real-time two-way communication

[1559] The device sends the questions or inquiries entered by the user to the server, which then uses its AI personality to generate an appropriate response and sends it back to the device, which then displays it to the user. This process is repeated, enabling real-time two-way communication.

[1560] 9. Saving conversation logs

[1561] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1562] 10. Analyzing conversation logs and improving response patterns

[1563] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1564] Specific examples

[1565] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what is bothering her at work today," the server generates Ken's response and provides appropriate advice. This conversation log is saved on the server and used for future conversations.

[1566] As described above, the present invention allows users to communicate in real time with the interlocutor that best suits them at any time, providing an immersive and satisfying experience.

[1567] The processing flow will be explained below.

[1568] Step 1:

[1569] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1570] Step 2:

[1571] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1572] Step 3:

[1573] The terminal transmits the input text data to the server.

[1574] Step 4:

[1575] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1576] Step 5:

[1577] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1578] Step 6:

[1579] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[1580] Step 7:

[1581] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[1582] Step 8:

[1583] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1584] Step 9:

[1585] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[1586] Step 10:

[1587] The server generates a response from the AI ​​personality "Ken" and sends it back to the device. For example, the server generates logical advice as text.

[1588] Step 11:

[1589] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[1590] Step 12:

[1591] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[1592] Step 13:

[1593] The server periodically analyzes the dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, enabling it to provide more appropriate responses in subsequent dialogues.

[1594] Example 1

[1595] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1596] With conventional AI dialogue systems, it can be difficult for users to select a conversation partner who matches their desired characteristics. Furthermore, interactions during a dialogue are not effectively saved and analyzed, making it difficult to improve the system for the next dialogue. This leaves users with a lack of satisfaction and immersion in the system, leading to the need for personalized advice.

[1597] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1598] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved session log and improving response patterns, means for transmitting text data from the user terminal to the server, and means for the server to generate responses based on the extracted keywords, thereby enabling users to communicate in real time with the interlocutor best suited to them.

[1599] "User terminal" means a device (e.g., smartphone, tablet, or PC) used by a user to operate the interface.

[1600] A "server" is a computer system that performs various processes such as text analysis, characteristic extraction, AI personality generation, log storage, and analysis.

[1601] A "database" is a system for storing dialogue logs and training data.

[1602] The "means for accepting user input" is a means for providing an interface for the user to input the characteristics of the interlocutor he / she desires.

[1603] "Morphological analysis" is a method of breaking down input text data into morphemes (the smallest units of language) and analyzing their meaning.

[1604] "Characteristic extraction" is the process of extracting important keywords related to the characteristics of the interlocutor from the analyzed text data.

[1605] An "AI personality" is an artificial intelligence character with specific personality traits and response patterns.

[1606] An "interactive session" refers to a series of interactions between a user and a generated AI personality.

[1607] A "session log" is data that records all interactions during an interactive session.

[1608] "Response pattern" refers to the pattern of responses that an AI personality generates in response to a specific input.

[1609] "Training data" is a dataset used to improve an AI model using historical data such as dialogue logs.

[1610] "Natural language processing technology" is a computer technology for analyzing, understanding, and generating human language.

[1611] A "generative AI model" is an artificial intelligence model that is trained to generate appropriate responses based on user input.

[1612] A "morpheme" is the smallest unit of language, an independent word or its equivalent that has meaning.

[1613] This invention is a system for providing immersive communication to users. This system generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user.

[1614] The system consists of the following components:

[1615] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1616] Server: A computer system that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis)

[1617] Database: A system for storing dialogue logs and training data

[1618] In this system, a user launches the application on a user device such as a smartphone, tablet, or PC. The device displays the question, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, if the user inputs "someone who is quiet and gives logical advice," the device sends this text data to the server.

[1619] The server performs morphological analysis on the received text data to extract important keywords such as "quiet," "logical," and "advice." Natural language processing technology (e.g., NLTK, SpaCy, etc.) is used for morphological analysis. The server then generates an AI personality based on the analyzed data. Specifically, it uses a generative AI model (e.g., GPT-3.5) to select and customize an AI model with response patterns that match specific personality traits.

[1620] The server creates a list of the generated AI personalities and sends it to the device. The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen, and the user selects the desired interlocutor. Once the dialogue session begins, the server and device work together to enable real-time two-way communication. The device sends the questions and consultation details entered by the user to the server, and the server uses the AI ​​personality to generate an appropriate response and send it back to the device. By repeating this process, a natural conversation progresses.

[1621] All messages and exchanges during a conversation session are stored in a database by the server as a conversation log. The server periodically analyzes the saved conversation log and uses it as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses being provided in subsequent conversations.

[1622] Specific examples

[1623] For example, a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics. This input data is sent to the server, and morphological analysis is used to extract the keywords "quiet," "logical," and "advice." Based on this, the server generates the optimal AI personality, "Ken." The device presents Erica with the options of "Ken" and other interlocutors, and Erica selects "Ken." When the dialogue session begins and Erica asks about "what's bothering her at work today," the server generates Ken's response and provides appropriate advice. This dialogue log is saved on the server and used for future dialogues.

[1624] Prompt Sentence Examples

[1625] User: Who do you want to talk to today?

[1626] AI: A quiet, logical adviser

[1627] User: Ken (quiet and logical)

[1628] AI: Hi Erica. What can we help you with today?

[1629] User: I'm having trouble with something at work today.

[1630] AI: What specifically are you worried about?

[1631] This system provides users with the most suitable interlocutor and enables real-time communication, thereby providing users with a high level of satisfaction and immersion.

[1632] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1633] Explanation of the system program processing flow and each processing step

[1634] Step 1:

[1635] A user operates a user device (smartphone, tablet, PC) to launch an application. The input is the user's action (tapping the app icon), and the output is the application's home screen displayed on the device. The operation is that the user device displays the application's logo for a few seconds, and then displays the home screen.

[1636] Step 2:

[1637] The terminal displays a question form to the user asking, "Who do you want to talk to today?" The input is a question template in the system, and the output is the question form displayed to the user. Specifically, the terminal displays a form containing a text box and a submit button on the screen.

[1638] Step 3:

[1639] The user inputs the desired characteristics of the interlocutor. For example, the user might input "someone who is quiet and gives logical advice." The input is the user's text data, and the output is the text data sent to the terminal by pressing the send button. The action is for the user to input the characteristics in the text box and click the send button.

[1640] Step 4:

[1641] The terminal sends the input text data to the server. The input is the user's text data, and the output is the text data sent to the server. In operation, the terminal encodes the user data into a packet format and sends it to the server via the Internet.

[1642] Step 5:

[1643] The server performs morphological analysis on the received text data and extracts important keywords such as "quiet," "logical," and "advice." The input is the text data sent from the device, and the output is the extracted keywords. In operation, the server analyzes the text data using natural language processing technology (e.g., NLTK, SpaCy, etc.) and breaks it down into morphemes.

[1644] Step 6:

[1645] The server generates an appropriate AI personality based on the extracted keywords. The input is the keywords extracted through analysis, and the output is an AI personality generated by a generative AI model (e.g., GPT-3.5). Specifically, the server uses the trained generative AI model to customize an AI model with response patterns that match specific personality traits.

[1646] Step 7:

[1647] The server creates a list of multiple AI personalities and sends it to the terminal. The input is a list of the generated AI personalities, and the output is the list sent to the user's terminal. In operation, the server packages the information about the AI ​​personalities into a list and sends it to the terminal.

[1648] Step 8:

[1649] The device displays candidates such as "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)" on the screen. The input is a list of AI personalities sent from the server, and the output is a list of candidates displayed on the screen. Specifically, the device displays a pop-up window and presents multiple AI personalities as options.

[1650] Step 9:

[1651] The user selects the desired AI personality (e.g., Ken). The input is the user's selection action, and the output is information about the selected AI personality. The user taps the desired interlocutor from the options on the screen.

[1652] Step 10:

[1653] The device notifies the server of the user's selection, and the server initializes an interactive session with the selected AI personality. The input is the user's selection data, and the output is the initial settings for the interactive session. Specifically, the device sends the selection data in packet form to the server, and the server sets up the interactive session.

[1654] Step 11:

[1655] The terminal displays a dialogue window, displaying a first-time greeting and question from the AI ​​interlocutor. The input is the initial response data from the server, and the output is a chat window and greeting message displayed to the user. Specifically, the terminal opens a chat window and displays a message such as "Hello, what would you like to discuss with us today?"

[1656] Step 12:

[1657] The user enters their question or inquiry into the terminal, which then sends it to the server. The input is the user's text data, and the output is the data to be sent to the server. The operation is that the user enters a question in the chat window and clicks the send button.

[1658] Step 13:

[1659] The server generates an appropriate response based on the data sent from the user's device and sends it back to the device. The input is the user's consultation data, and the output is the response generated by the server. Specifically, the server uses a generative AI model to create an appropriate response and sends it to the user's device.

[1660] Step 14:

[1661] The terminal displays the response sent from the server to the user. The input is the response data from the server, and the output is a text message displayed to the user. In operation, the terminal displays the response from the server in a chat window. This process is repeated, achieving real-time two-way communication.

[1662] Step 15:

[1663] The server stores all messages and exchanges during an interactive session as an interaction log in a database. The input is all exchanges during the interaction, and the output is the interaction log stored in the database. Specifically, the server collects all messages in real time and stores them in the database.

[1664] Step 16:

[1665] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the response patterns of the AI ​​personality. The input is the stored dialogue logs, and the output is the improved response patterns. Specifically, the server analyzes the dialogue logs and uses them as a training dataset for the generative AI model. The improved model is reflected in the next session.

[1666] (Application example 1)

[1667] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1668] Today's consumers increasingly desire the support of virtual assistants that can provide specific advice and recommendations, helping them make faster and more accurate product selections. However, many existing virtual assistants are unable to fully address the individual needs and preferences of users and often only provide monotonous responses. This has led to a demand for the development of virtual assistant systems that can provide more personalized assistance to users.

[1669] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1670] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, and means for generating an AI personality as a shopping assistant tailored to the user's individual preferences when the user accesses a virtual store and supporting the user with shopping. This allows the user to proceed with shopping while interacting in real time with a shopping assistant tailored to the user's individual preferences.

[1671] "Means for accepting user input" refers to devices or software that have an interface or function for inputting the characteristics and requests of the desired interlocutor into the system.

[1672] "Means for morphologically analyzing text data entered by a user" refers to natural language processing technology for analyzing text data entered by a user and identifying its components.

[1673] The "means for extracting personality traits from analyzed text data" refers to a function or algorithm for extracting the traits of the interlocutor that the user desires from the data obtained by morphological analysis.

[1674] "Means for generating an AI personality based on extracted personality traits" refers to a mechanism or software that generates an AI personality with specific characteristics based on analyzed information.

[1675] "Means for presenting the generated AI personality to the user" refers to an interface or function that displays options for the generated AI personality to the user and encourages them to make a selection.

[1676] "Means for initiating an interactive session with a user-selected AI personality" refers to the functionality or protocols for initiating two-way communication with a user-selected AI personality.

[1677] The "means for storing interactions during an interactive session" refers to a technique for recording all interactions that occur during an interactive session and storing them for later reference.

[1678] "Means for analyzing stored dialogue logs and improving response patterns" refers to functions and algorithms that analyze stored dialogue logs and use the information obtained from them to improve the AI's response patterns.

[1679] "Means of generating an AI personality as a shopping assistant tailored to individual preferences when a user accesses a virtual store and supporting the user in their shopping" refers to functions and software that automatically generate a shopping assistant tailored to the characteristics and requests of the user when the user uses a virtual store, and support the shopping process.

[1680] This invention is a system that provides immersive communication to users, and in particular, demonstrates its application as a shopping assistant in a virtual store. This system uses specific hardware and software to generate an AI personality based on the characteristics of the person the user desires to interact with, and to realize real-time two-way communication.

[1681] System configuration

[1682] User terminal

[1683] Hardware: Smartphones, tablets, computers, etc.

[1684] Software: Web browser or dedicated application

[1685] server

[1686] Hardware: High-performance server machine

[1687] software:

[1688] Natural Language Processing (NLP) libraries (e.g., spaCy)

[1689] Generative AI models (e.g., GPT-2 by OpenAI)

[1690] Database (e.g. PostgreSQL)

[1691] Web server (e.g. Nginx, Gunicorn)

[1692] Processing flow

[1693] 1. User Input

[1694] The user starts a dedicated application using the user terminal. The application displays a question to the user: "What kind of person do you want to talk to today?" The user then inputs the characteristics of the person they want to talk to. For example, they might input "someone who is knowledgeable about fashion and can tell me about trends."

[1695] 2. Sending and analyzing input data

[1696] The device sends the entered text data to a server, which then receives the data and performs morphological analysis using an NLP library to extract important keywords such as "fashion" and "trend."

[1697] 3. AI personality generation

[1698] The server generates an appropriate AI personality using a generative AI model (e.g., GPT-2) based on the extracted keywords. A prompt such as "You are a shopping assistant with a fashion and trend-oriented personality. Please provide the user with the best advice." is used.

[1699] 4. Starting an interactive session

[1700] The generated AI personalities are presented to the user as multiple options. A dialogue session is initiated with the AI ​​personality selected by the user, and real-time two-way communication takes place. Questions and inquiries entered by the user are sent to the server, which generates an appropriate response and returns it to the user. This process is repeated.

[1701] 5. Saving and analyzing conversation logs

[1702] The server stores logs of all interaction sessions in a database, which are periodically analyzed and used as training data to provide better responses in subsequent interactions.

[1703] Specific examples

[1704] For example, a user named Erika launches a dedicated application and enters "someone who is knowledgeable about fashion and can tell me about trends" as the desired characteristic of her interlocutor. The server analyzes this input data and generates the optimal AI personality. The device presents the generated AI personality to Erika, who selects "Fashion Assistant." When the dialogue session begins and Erika asks, "What are the trends this season?", the server generates a response for the AI ​​assistant and provides appropriate trend information. This dialogue log is saved on the server and used for future dialogues.

[1705] An example prompt is:

[1706] You are a shopping assistant with a fashion and trend-oriented nature. Give your customers the best advice.

[1707] As a result of the above, this invention allows users to proceed with their shopping while interacting in real time with a shopping assistant that is tailored to their individual preferences, resulting in a more satisfying shopping experience.

[1708] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1709] Step 1:

[1710] Input: The user starts the dedicated application using the user terminal.

[1711] What it does: The device displays the application home screen and a question form asking, "Who do you want to talk to today?"

[1712] Output: The user inputs the desired characteristics of the interlocutor.

[1713] Step 2:

[1714] Input: The user inputs, for example, "someone who is knowledgeable about fashion and can tell me about trends" as the characteristics of the person he or she desires to talk to.

[1715] Operation: The terminal sends the entered text data to the server.

[1716] Output: The input data is sent to the server.

[1717] Step 3:

[1718] Input: Text data received by the server

[1719] How it works: The server uses a natural language processing (NLP) library (e.g., spaCy) to perform morphological analysis and extract keywords such as "fashion" and "trend."

[1720] Output: Keywords extracted as analysis results

[1721] Step 4:

[1722] Input: Keywords extracted as analysis results

[1723] How it works: The server uses a generative AI model (e.g., GPT-2) to generate an appropriate AI personality. The prompt used is, "You are a shopping assistant with a fashion and trending personality. Please provide the user with the best advice."

[1724] Output: Generated AI personality

[1725] Step 5:

[1726] Input: Generated AI personality

[1727] How it works: The server creates a list of the generated AI personalities as multiple options and sends it to the device.

[1728] Output: AI personality options displayed on the device

[1729] Step 6:

[1730] Input: AI personality options displayed on the device

[1731] How it works: The user selects the AI ​​personality they want, for example "Fashion Assistant."

[1732] Output: User-selected AI personality

[1733] Step 7:

[1734] Input: User-selected AI personality

[1735] Operation: The device starts an interactive session with the selected AI personality. The server prepares the necessary initial data and initializes the interactive session.

[1736] Output: Interactive session started

[1737] Step 8:

[1738] Input: User's question or inquiry

[1739] How it works: The device sends the question or inquiry entered by the user to the server, which then generates an appropriate response. This response is generated using a generative AI model. For example, if a user asks, "What's trending this season?", the server generates a response and sends it back to the device.

[1740] Output: The AI ​​assistant's response that is displayed to the user

[1741] Step 9:

[1742] Input: All messages and exchanges during an interactive session

[1743] How it works: The server stores all messages and exchanges from an interactive session in a database.

[1744] Output: Saved interaction logs

[1745] Step 10:

[1746] Input: Saved conversation log

[1747] How it works: The server periodically analyzes the stored dialogue logs and uses them as training data to refine the AI ​​personality's response patterns, which will result in more appropriate responses being provided in future dialogues.

[1748] Output: Improved response pattern

[1749] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1750] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses. Below, as an embodiment of the present invention, the processing of the system's program is described in detail in natural language.

[1751] Basic system configuration

[1752] The system consists of the following elements:

[1753] User terminal: The device through which the user operates the interface (e.g., smartphone, tablet, PC)

[1754] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[1755] Database: A database that stores interaction logs and training data.

[1756] Emotion Engine: Technology for analyzing and recognizing user emotions

[1757] System processing flow

[1758] 1. The user launches the application

[1759] The user taps the application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1760] 2. Accept user input

[1761] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1762] 3. Sending input data

[1763] The terminal transmits the input text data to the server.

[1764] 4. Morphological analysis and feature extraction

[1765] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1766] 5. AI personality generation

[1767] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1768] 6. User Sentiment Analysis

[1769] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1770] 7. Present interlocutor options

[1771] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1772] 8. Starting an interactive session

[1773] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1774] 9. Real-time two-way communication

[1775] The device sends the user's input questions or consultation details (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1776] 10. Saving conversation logs

[1777] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1778] 11. Analyzing conversation logs and improving response patterns

[1779] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1780] Specific examples

[1781] For example, if a user named Erica starts the application and enters "someone who is quiet and gives logical advice" as the desired trait of a conversation partner, the server analyzes this input data and generates the optimal AI personality, "Ken." The device presents Erica with the option of "Ken" and other conversation partners, and Erica selects "Ken." When the conversation session begins and Erica asks about "what's bothering her at work today," the emotion engine analyzes Erica's emotion as "anxiety," and the server generates Ken's response. Ken offers kind words and logical solutions to alleviate Erica's anxiety. This conversation log is saved on the server and used for future conversations.

[1782] As described above, the present invention allows a user to communicate in real time with the person who best suits him or her at any time, and to obtain a satisfying experience that is tailored to his or her emotions.

[1783] The processing flow will be explained below.

[1784] Step 1:

[1785] A user taps an application icon on their smartphone to launch the application, and the device displays the application's home screen.

[1786] Step 2:

[1787] The terminal displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1788] Step 3:

[1789] The terminal transmits the input text data to the server.

[1790] Step 4:

[1791] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1792] Step 5:

[1793] The server generates an AI personality with the corresponding personality traits based on the extracted keywords, specifically selecting and customizing an AI model with a calm and logical response pattern.

[1794] Step 6:

[1795] The emotion engine installed on the server analyzes emotions from the text entered by the user, determining emotions such as "anxiety," "joy," and "anger."

[1796] Step 7:

[1797] The server creates a list of multiple AI personalities (e.g., "Ken (quiet and logical)" and "Kazuki (emotional and encouraging)") and sends this to the device.

[1798] Step 8:

[1799] The terminal displays a list of interlocutor options to the user, and the user selects the desired interlocutor, for example, the user selects "Ken."

[1800] Step 9:

[1801] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1802] Step 10:

[1803] The terminal sends the questions or consultation content entered by the user (for example, "What are you worried about at work today?") to the server.

[1804] Step 11:

[1805] The server considers the user's emotions analyzed by the emotion engine and generates a response from the AI ​​personality "Ken." For example, if the user's emotion is recognized as "anxiety," Ken will provide reassuring advice in a calm tone.

[1806] Step 12:

[1807] The device receives a response from the server and displays it to the user. This process is repeated in real time, allowing for continuous two-way communication.

[1808] Step 13:

[1809] The server stores all messages and exchanges from the interactive session in a database as an interaction log.

[1810] Step 14:

[1811] The server periodically analyzes the saved dialogue logs and uses them as training data to improve the AI ​​personality "Ken's" response patterns, providing more appropriate responses in subsequent dialogues.

[1812] Example 2

[1813] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1814] Conventional dialogue systems can only provide fixed responses to user input, making it difficult to generate responses that take the user's emotions and circumstances into account. This has led to the issue of not being able to provide a satisfying dialogue experience for users. Furthermore, the learning function required to appropriately utilize dialogue logs and generate more appropriate responses for the next dialogue has also been insufficient.

[1815] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1816] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data input by the user, means for extracting characteristics from the analyzed text data, means for generating an AI personality based on the extracted characteristics, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, and means for generating responses based on the analyzed emotion information. This makes it possible to generate appropriate responses in real time according to the user's emotions and situation, and to provide more appropriate responses in the next interaction by utilizing the interaction log.

[1817] "User" means any person or entity that interacts with the System.

[1818] "Means for accepting input" refers to an interface that allows a user to input information into the system, and is a system that includes input devices such as a keyboard, touch screen, or voice input.

[1819] "Morphological analysis" is a technique that divides text into words and morphemes and analyzes the meaning and grammatical information of each word.

[1820] The "means for extracting characteristics" is a function for identifying and extracting important keywords and characteristics from morphologically analyzed text data.

[1821] An "artificial intelligence personality" is a virtual conversational agent generated based on extracted characteristics, capable of engaging in natural conversations with users.

[1822] "Means for presentation" refers to a method for visually or audibly displaying the personality of the generated AI to the user, and includes screen display, audio announcement, etc.

[1823] "Interaction session" refers to a series of interactions between a user and a selected AI personality, including multiple message exchanges.

[1824] "Means for storing interactions" refers to a function that records messages and information exchanged between the user and the AI ​​during an interaction session and stores them for later use.

[1825] An "interaction log" is a record of all messages and actions exchanged during an interaction session.

[1826] "Means for improving response patterns" is a function that analyzes saved dialogue logs and learns and improves the AI's responses so that it can respond more appropriately in the next dialogue.

[1827] "Means for analyzing emotions" refers to technology for reading and identifying emotions (e.g., joy, anger, sadness, etc.) from the user's input text or voice.

[1828] The "means for generating a response" is a function that allows the AI ​​personality to create an appropriate response based on the analyzed emotional information and extracted characteristics.

[1829] The present invention is a system that provides immersive communication to users and combines it with an emotion engine that recognizes the user's emotions. This system not only generates an AI personality tailored to the user's needs and realizes two-way dialogue with the user, but also analyzes the user's emotions and adjusts responses accordingly.

[1830] The basic system configuration is as follows:

[1831] Device: The device through which the user operates the interface (e.g., smartphone, tablet, PC).

[1832] Server: A server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, and sentiment analysis).

[1833] Database: A database that stores dialogue logs and training data.

[1834] Emotion engine: Technology for analyzing and recognizing user emotions.

[1835] The server analyzes the text data entered by the user using a morphological analysis tool called "MeCab," for example, and extracts important keywords. Based on these keywords, the server generates an AI personality using a generative AI model (e.g., GPT-3). Furthermore, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions and adjusts responses based on the analysis results.

[1836] As a concrete example, a user named Erica launches the application and inputs the desired characteristics of a partner: "Someone who is quiet and gives logical advice." The server that receives this input performs morphological analysis and generates an AI personality called "Ken" using GPT-3 based on the extracted keywords. The device then presents "Ken" and other partner options to Erica.

[1837] When Erica selects "Ken" and a dialogue session begins, the server uses the emotion engine to analyze Erica's input and determine her emotions. For example, if Erica asks about her worries at work today, the emotion engine analyzes her emotion as "anxiety." In response, the server generates a response from "Ken" that offers kind words and logical solutions to alleviate Erica's anxiety.

[1838] This dialogue log is stored in a database and analyzed by the server, which uses it as learning data to provide more appropriate responses in subsequent dialogues.

[1839] An example prompt is, "When Erica is feeling anxious about work today, how would Ken, the calm and logical AI personality, respond?"

[1840] In this way, the present invention allows users to communicate with the person who best suits them in real time at any time, providing a satisfying experience tailored to their emotions.

[1841] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1842] Step 1:

[1843] The user launches an application.

[1844] A user taps an application icon on their smartphone to launch the application. The input is the user's action (selecting and tapping the application). The output is the device displaying the application's home screen, which includes the message "Hello! Who would you like to talk to today?"

[1845] Step 2:

[1846] The device displays the input form.

[1847] The device displays a question form to the user asking, "What kind of person do you want to talk to today?" The required input information is the user's action (launching the app) that triggers the display of the form. The output is a text box where the user can enter the characteristics of the person they want to talk to. For example, the user might enter, "Someone who is quiet and gives logical advice."

[1848] Step 3:

[1849] The terminal transmits the text data entered by the user to the server.

[1850] The terminal sends the previously entered text data to the server. The input data is the characteristic information entered by the user. As an output, the sent data arrives at the server and the next processing step is initiated.

[1851] Step 4:

[1852] The server performs morphological analysis on the received text data.

[1853] The server performs morphological analysis of the input text data using an analysis tool (e.g., "MeCab"). The input text data is divided into words and morphemes, and the meaning and grammatical information of each word is analyzed. Important keywords such as "quiet," "logical," and "advice" are extracted as output.

[1854] Step 5:

[1855] The server generates an AI personality based on the data after characteristic extraction.

[1856] The server generates an AI personality using a generative AI model (e.g., GPT-3) based on the extracted keywords. The input is the characteristic keywords extracted through morphological analysis. The output is, for example, "A quiet and logical AI personality named 'Ken.'"

[1857] Step 6:

[1858] The server analyzes the user's emotions.

[1859] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from the user's input text or voice. The input data is the user's text or voice, and the analysis results in emotion labels such as "anger," "joy," or "sadness." In this example, the user's text, "Work wasn't going well today," is analyzed as "anxiety."

[1860] Step 7:

[1861] The server generates multiple AI personalities and sends them to the device.

[1862] The server sends the generated AI personality (e.g., "Ken (quiet and logical)" or "Akira (emotional and encouraging)") to the terminal, which then presents it to the user. The terminal has data for each generated AI personality as input, and displays options to the user as output. For example, the options "Ken" and "Akira" are displayed to the user.

[1863] Step 8:

[1864] The terminal notifies the server of the user's selection, initiating an interactive session.

[1865] The user selects the desired interlocutor, and the terminal notifies the server of the selection. The input is the user's selection action. The output is the initialization of the interaction session, and the terminal displays an initial greeting or question. For example, the message "Hello, I'm Ken! How's your day?" is displayed on the terminal.

[1866] Step 9:

[1867] The user inputs a question or inquiry, and the server generates a response accordingly.

[1868] The user inputs the specific content of their problem (e.g., "What are you worried about at work today?"), and the device sends that data to the server. The input data is the user's problem content and an emotion label. The server receives this and generates an appropriate response based on the emotion analysis results. For example, the output response might be, "That's tough. Can you tell me specifically what happened?"

[1869] Step 10:

[1870] The server logs the interactive session.

[1871] The server stores all messages and exchanges during a conversation session in a database as a conversation log. The input data are messages and emotion labels during the conversation session. The output is the stored log entries that will be used to refine the next response pattern.

[1872] Step 11:

[1873] The server analyzes the dialogue logs and refines response patterns.

[1874] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns. The input data is the stored dialogue logs, and the output is improved response patterns to improve user satisfaction in subsequent dialogue sessions. For example, based on Erica's dialogue history, the system can learn her preferences and tendencies and provide more accurate advice.

[1875] (Application example 2)

[1876] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1877] Conventional dialogue systems have difficulty generating appropriate responses based on the user's emotions, making it impossible to provide dialogue that meets individual needs. Furthermore, when interacting with robots in factories or workplaces, there is a need to provide advice and support that takes into account the worker's emotions. In this situation, it is necessary to develop a system that incorporates emotion analysis.

[1878] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1879] In this invention, the server includes means for accepting user input, means for morphologically analyzing text data entered by the user, means for extracting personality traits from the analyzed text data, means for generating an AI personality based on the extracted personality traits, means for presenting the generated AI personality to the user, means for starting an interaction session with the AI ​​personality selected by the user, means for saving exchanges during the interaction session, means for analyzing the saved interaction log and improving response patterns, means for analyzing the user's emotions, means for generating responses according to the analyzed emotions, and means for presenting the responses to the user via a support device. This enables appropriate interaction that takes emotions into consideration.

[1880] "User" refers to the person who operates and interacts with the system.

[1881] "Input" refers to the text data and voice data that the user provides to the system.

[1882] "Means" refers to a method or apparatus for performing a particular function within a system.

[1883] "Morphological analysis" refers to the process of breaking down text data into words and analyzing them.

[1884] "Personality traits" refer to specific personality traits necessary for generating an AI personality.

[1885] "AI personality" refers to a personality model created by artificial intelligence that can converse with users.

[1886] A "session" refers to a series of conversations between a user and an AI personality.

[1887] "Interaction Log" means a record of all messages exchanged during an interaction session.

[1888] "Analysis" refers to the process of finding new information and areas for improvement based on stored data.

[1889] "Emotion" refers to the psychological state that a user exhibits during a conversation.

[1890] "Analysis" refers to the process of examining data in detail to understand its contents.

[1891] "Response" refers to the answer or reaction that the system gives back to the user.

[1892] "Support device" refers to a device that presents the output of the system to the user.

[1893] "Presentation" refers to the act of displaying system-generated information or responses to the user.

[1894] This invention is a system that provides immersive communication to users and can recognize their emotions and adjust responses accordingly. This system generates an AI personality based on the user's needs and engages in real-time dialogue.

[1895] Basic system configuration

[1896] The system consists of the following elements:

[1897] User terminal: The device through which the user operates the interface (e.g., smart glasses, smartphone, computer)

[1898] Server: Server that performs various processes (text analysis, characteristic extraction, AI personality generation, log storage, analysis, emotion analysis)

[1899] Database: A database that stores interaction logs and training data.

[1900] Emotion Engine: Technology for analyzing and recognizing user emotions

[1901] Hardware and software used

[1902] Hardware: smart glasses (e.g. Microsoft HoloLens), smartphone, PC

[1903] Software: Python, transformers library, sentiment analysis model (EmoRoBERTa), text generation model (GPT-3.5-turbo)

[1904] Processing flow

[1905] 1. The user launches the application

[1906] The user taps the application icon on their smart glasses or smartphone to launch the application, and the device displays the application's home screen.

[1907] 2. Accept user input

[1908] The device displays a question form to the user asking, "What kind of person do you want to talk to today?", and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1909] 3. Sending input data

[1910] The terminal transmits the input text data to the server.

[1911] 4. Morphological analysis and feature extraction

[1912] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1913] 5. AI personality generation

[1914] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it selects and customizes an AI model with a calm and logical response pattern.

[1915] 6. User sentiment analysis

[1916] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1917] 7. Present interlocutor options

[1918] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1919] 8. Starting an interactive session

[1920] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1921] 9. Real-time two-way communication

[1922] The device sends the question or consultation content entered by the user (for example, "What are you worried about at work today?") to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[1923] 10. Saving conversation logs

[1924] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[1925] 11. Analyzing conversation logs and improving response patterns

[1926] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[1927] Specific examples

[1928] For example, if a user launches the application and enters "someone who is quiet and gives logical advice" as the desired interlocutor's characteristics, the server will analyze this input data and generate the optimal AI personality. The device will present the user with several options, and the user will select one. When the dialogue session begins and the user complains that "the machine is not working properly," the emotion engine will interpret this as "anxiety," and the server will generate a response accordingly ("Don't worry, check the instructions in the operating manual and try restarting it as directed") and display it on the smart glasses.

[1929] Prompt Sentence Examples

[1930] Example prompt: "The worker is feeling anxious about the current task. Please consider his / her feelings and provide appropriate advice."

[1931] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1932] Step 1:

[1933] The user launches the application.

[1934] Specific behavior:

[1935] The user turns on the smart glasses or smartphone and taps the application icon.

[1936] input:

[1937] Tap the application icon

[1938] output:

[1939] The application home screen will appear on your device.

[1940] Step 2:

[1941] Accepts user input.

[1942] Specific behavior:

[1943] The device displays a form asking the user, "What kind of person do you want to talk to today?" and the user inputs the characteristics of the person they want to talk to. For example, they might input, "Someone who is quiet and gives logical advice."

[1944] input:

[1945] Characteristics of desired interlocutor input by user

[1946] output:

[1947] The entered characteristics are recorded as text data.

[1948] Step 3:

[1949] Sending input data.

[1950] Specific behavior:

[1951] The terminal transmits the input text data to the server.

[1952] input:

[1953] Text data entered by the user

[1954] output:

[1955] The text data is sent to the server.

[1956] Step 4:

[1957] Morphological analysis and feature extraction.

[1958] Specific behavior:

[1959] The server performs morphological analysis on the received text data, splitting and analyzing the words in the text to extract important keywords (such as "quiet," "logical," and "advice").

[1960] input:

[1961] Text data sent to the server

[1962] output:

[1963] Important keywords are extracted.

[1964] Step 5:

[1965] AI personality generation.

[1966] Specific behavior:

[1967] The server generates an AI personality with the corresponding personality traits based on the extracted keywords. Specifically, it uses a generative AI model to select and customize an AI personality with a calm and logical response pattern.

[1968] input:

[1969] Extracted keywords

[1970] output:

[1971] Generated AI personality

[1972] Step 6:

[1973] User sentiment analysis.

[1974] Specific behavior:

[1975] The emotion engine installed on the server analyzes emotions from the user's input text or voice, outputting emotion labels such as "anger," "joy," and "sadness."

[1976] input:

[1977] User input text or voice

[1978] output:

[1979] Emotion Labels

[1980] Step 7:

[1981] Presenting interlocutor options.

[1982] Specific behavior:

[1983] The server creates a list of multiple AI personalities (e.g., "quiet personality" or "logical personality") and sends it to the device. The device displays the list to the user, who can then select the desired interlocutor.

[1984] input:

[1985] List of generated AI personalities

[1986] output:

[1987] The list of contacts presented to the user

[1988] Step 8:

[1989] Start an interactive session.

[1990] Specific behavior:

[1991] The device notifies the server of the user's selection, and the server initiates a dialogue session with the selected AI personality. The device displays a dialogue window, displaying an initial greeting and question.

[1992] input:

[1993] Notification of the user's choice of AI personality

[1994] output:

[1995] Displaying a dialogue window, initial greetings and questions

[1996] Step 9:

[1997] Real-time two-way communication.

[1998] Specific behavior:

[1999] The device sends the questions and consultation details entered by the user to the server. The server generates a response that takes the user's emotions into consideration and provides appropriate advice. For example, if the user is feeling "anger," the AI ​​personality will respond in a calm tone.

[2000] input:

[2001] User questions, consultation details, and emotion labels

[2002] output:

[2003] Emotion-aware responses

[2004] Step 10:

[2005] Saving conversation logs.

[2006] Specific behavior:

[2007] The server stores all messages and exchanges of the interactive session in a database as an interaction log.

[2008] input:

[2009] Interactive session messages and exchanges

[2010] output:

[2011] Interaction logs stored in a database

[2012] Step 11:

[2013] Analyzing dialogue logs and refining response patterns.

[2014] Specific behavior:

[2015] The server periodically analyzes the stored dialogue logs and uses them as training data to improve the AI ​​personality's response patterns, which will result in more appropriate responses in subsequent dialogues.

[2016] input:

[2017] Saved conversation logs

[2018] output:

[2019] Improved response patterns

[2020] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2021] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2022] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2023] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2024] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2025] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2026] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2027] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2028] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2029] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2030] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2031] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2032] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2033] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2034] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2035] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2036] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2037] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system t...

Claims

1. means for accepting user input; A means for performing morphological analysis on text data input by a user; means for extracting personality traits from the analyzed text data; A means for generating an AI personality based on the extracted personality traits; and A means for presenting the generated AI personality to the user; a means for initiating an interactive session with the user's selected AI personality; a means for storing communications during an interactive session; A means for analyzing the stored dialogue logs and improving response patterns; A system including:

2. 10. The system of claim 1, further comprising an interface for a user to input desired interlocutor characteristics.

3. 10. The system of claim 1, wherein the dialogue log is analyzed and used as training data to refine response patterns for subsequent dialogues.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A