system

The system addresses the lack of realistic conversational experiences in language learning by allowing data upload, management, and model training, facilitating simulations with celebrities and professionals to improve learner motivation and effectiveness.

JP2026022424APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123941
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Conventional language learning methods fail to provide realistic conversational experiences with real people, limiting opportunities for learners to simulate conversations with celebrities or professionals, and lack a system to reward data providers sustainably.

Method used

A system that allows conversation data providers to upload data, which is stored, managed, and used to train large-scale language models, enabling realistic simulations and interaction logs for improved learning, with the server generating responses based on user inputs.

Benefits of technology

Enables users to engage in realistic conversation simulations with celebrities and professionals, enhancing learning motivation and effectiveness by providing interactive and engaging content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022424000001_ABST
    Figure 2026022424000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes means for a conversation data provider to upload conversation data from a terminal, means for a server to store the received conversation data in a database, means for the server to supply the stored conversation data to an LLM (Large Language Model) for training, means for the server to store a trained conversation model in a repository, means for a user to select a conversation partner and a situation from the terminal and request a simulation, means for the server to call the conversation model in response to the user's request, generate a response, and transmit the response to the terminal, and means for the server to store interactions between the user and the conversation model as a log.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional language learning methods have made it difficult for foreign language learners to gain realistic conversational experience with real people. In particular, opportunities to simulate conversations with celebrities or professionals have been limited, preventing learners from maximizing their motivation and learning effectiveness. Furthermore, the lack of a system to provide appropriate rewards to those who provide conversation data has made it difficult to provide sustainable conversation data. The present invention aims to solve these problems. [Means for solving the problem]

[0005] The system according to the present invention includes a means for a conversation data provider to upload conversation data from a terminal, thereby collecting a variety of conversation data. The server further includes a means for storing the received conversation data in a database, enabling data management and indexing. The stored conversation data is then supplied to a large-scale language model (LLM) by the server for training. The system also includes a means for storing the conversation model generated by the training means in a repository, and when a user selects a conversation partner and situation from the terminal and requests a simulation, the server invokes the appropriate model, generates a response, and sends it to the terminal. The server also includes a means for storing interactions between the user and the conversation model as a log, which can be used as subsequent training data to improve the quality of the conversation. Thus, the present invention provides a realistic conversation simulation and enables users to learn.

[0006] "Conversation data provider" refers to any individual or organization that provides conversation data to the system, including celebrities and professionals.

[0007] "Terminal" refers to an electronic device, such as a computer, smartphone, or tablet, used by a provider or user to access the system.

[0008] "Conversation data" refers to the content of a conversation recorded in audio or text format, and is used as learning data or simulation data.

[0009] "Index" refers to the index created by the server to efficiently search and manage saved conversation data.

[0010] "Server" refers to a central management unit for receiving, storing, processing, managing and distributing conversation data sent from terminals.

[0011] "Database" refers to a system for centrally storing and managing received conversation data and generated models.

[0012] "LLM (Large-scale Language Model)" refers to a machine learning model that enables natural language generation and conversational understanding by training on large amounts of text data.

[0013] "Repository" refers to a storage system for storing trained conversational models so that they can be retrieved as needed.

[0014] "User" refers to an individual who uses the system to conduct conversation simulations, such as a language learner.

[0015] "Simulation" refers to the process of recreating a realistic conversational experience based on the conversation partner and situation selected by the user.

[0016] "Response" refers to the natural language reply generated by the conversation model in response to user input.

[0017] "Logs" refer to data that records interactions between users and conversation models, and are also used as subsequent training data. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is a system that allows users to learn a language through realistic conversation simulations. This system provides simulations based on conversation data with famous people and professionals, and aims to improve learners' motivation and learning effectiveness.

[0040] System Configuration

[0041] The system includes the following major components:

[0042] 1. Terminal

[0043] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[0044] An interface for conversation data providers to upload conversation data.

[0045] An interface for users to request conversation simulations and view the simulation results.

[0046] 2. Server

[0047] The ability to receive, save, clean, and store conversation data in a database.

[0048] A function that trains LLMs (large-scale language models) based on saved conversation data.

[0049] Ability to save trained conversational models and recall them whenever needed.

[0050] The ability to use conversational models to generate realistic responses based on user requests.

[0051] A function that saves conversation logs and uses them as training data for future use.

[0052] 3. Database

[0053] A system that centrally stores and manages received conversation data and generated conversation models.

[0054] Collection and handling of conversation data

[0055] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[0056] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[0057] Training the model

[0058] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[0059] User Interaction

[0060] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[0061] Terminal: The selection is sent to the server as a request.

[0062] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[0063] Real-time conversation simulation

[0064] User: Enter the following conversation phrase and send it to the server via the device.

[0065] Server: Analyzes the user's input data, generates the next response, and sends it to the device. This interaction occurs in real time, allowing the user to simulate a realistic conversation experience with a real person.

[0066] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[0067] Specific examples

[0068] A user simulates an interview with celebrity B.

[0069] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[0070] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[0071] 3. Terminal: The initial response is displayed to the user.

[0072] 4. User: Enter the following question and submit it to the server:

[0073] 5. Server: Generates the following response and sends it to the device:

[0074] 6. Server: Stores the conversation log.

[0075] In this way, users can advance their language learning while simulating a real-life interview experience with celebrity B.

[0076] The processing flow will be explained below.

[0077] Step 1: Upload your conversation data

[0078] Terminal: Conversation data providers use terminals to upload conversation data (audio files or text files) to the system.

[0079] Server: Stores the received conversation data in temporary storage and checks the data integrity.

[0080] Step 2: Storing and indexing conversation data

[0081] Server: Stores the verified conversation data in a database.

[0082] Server: Creates an index of the stored data to facilitate future search and management.

[0083] Step 3: Cleaning the conversation data

[0084] Server: Cleans the stored conversation data, removing noise and privacy information.

[0085] Step 4: Generate a training set

[0086] Server: Converts the cleaned data into a training set for LLMs (large-scale language models).

[0087] Step 5: Train the model

[0088] Server: Feeds the training set to the LLM, allowing it to learn specific person speaking styles and phrases.

[0089] Server: Stores the trained model in a repository.

[0090] Step 6: User login and profile verification

[0091] User: Log in to the system and check your profile information.

[0092] Terminal: Sends the user's login information to the server for authentication.

[0093] Step 7: Choose your conversation partner and situation

[0094] User: Select a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the terminal.

[0095] Terminal: Sends the selection to the server as a request.

[0096] Step 8: Generate an initial response

[0097] Server: Based on the user's request, it invokes the appropriate conversation model to generate the initial response.

[0098] Server: Sends the generated initial response to the terminal.

[0099] Step 9: View the initial response

[0100] Terminal: Displays to the user the initial response received from the server.

[0101] Step 10: Enter conversation phrases

[0102] User: Enter the following conversation phrase into your device:

[0103] Terminal: Sends user input data to the server.

[0104] Step 11: Generate and Send Response

[0105] Server: Parses the user's input data and generates the following response:

[0106] Server: Sends the generated response to the terminal.

[0107] Step 12: Save the conversation log

[0108] Server: Conversation logs with the user are stored in a database and used as training data for future use.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] Language learning is becoming increasingly important in modern times, and learning through real conversations is particularly effective. However, opportunities to experience real conversations are limited, making it difficult to learn through conversation simulations with specific celebrities or professionals. Furthermore, systems that generate responses in real time in response to user requests are not common, making it difficult to provide a conversation experience that is close to reality. This has led to problems such as insufficient improvement in user motivation and effectiveness of learning.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) and train it; a means for the server to store the trained conversation model in a repository; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response and send it to the terminal; a real-time conversation simulation means for the user to input the next conversation phrase and the server to generate a response and send it to the terminal; and a means for the server to store the interaction between the user and the conversation model as a log. This enables language learning through realistic conversation simulations with specific celebrities or professionals, thereby improving the motivation and effectiveness of users' learning.

[0114] "Conversation data provider" refers to a user or individual who provides conversation data to the conversation simulation system.

[0115] "Terminal" refers to a device, such as a computer, smartphone, or tablet, used by a conversation data provider or user to access the system.

[0116] "Conversation data" refers to materials such as audio files and text files used for conversation simulation.

[0117] "Server" refers to a computer system that receives, stores, processes data, trains models, etc.

[0118] "Database" refers to a system for efficiently storing, managing, and searching conversation data.

[0119] "LLM (Large-scale Language Model)" refers to an AI model for natural language processing that is trained based on massive amounts of conversational data.

[0120] "Training Set" refers to the collection of cleaned conversational data used to train the LLM.

[0121] "Repository" refers to a storage system for storing trained conversational models.

[0122] "User" refers to an individual who utilizes the conversation simulation system to simulate a real conversation experience.

[0123] "Situation" refers to the conversation format or setting selected by the user, such as an interview format.

[0124] A "request" refers to a request or instruction a user makes to a system.

[0125] "Response" means any reply or response content generated by LLM to a User's request.

[0126] "Real-time conversation simulation" refers to a process in which the server instantly generates and returns a response to a conversation phrase entered by the user.

[0127] "Conversation log" refers to data that records the interactions between a user and a conversation model.

[0128] "Cleaning" refers to the process of removing unnecessary noise and private information from conversation data.

[0129] The present invention is a system that allows users to learn a language through realistic conversation simulations. One feature of this system is that it provides simulations based on conversation data with famous people and professionals, thereby improving learners' motivation and learning effectiveness. An embodiment of the present invention is described in detail below.

[0130] This system mainly consists of three components: a terminal, a server, and a database.

[0131] Terminal

[0132] Terminals are devices through which conversation data providers and users access the system, and examples include computers, smartphones, and tablets. Conversation data providers use these terminals to access an interface for uploading conversation data such as audio files and text files to the system. Users can also request conversation simulations and view the simulation results through the terminals.

[0133] server

[0134] The server is responsible for the core processing of the system. First, it receives conversation data uploaded from the device and temporarily stores it in storage. Next, it checks the integrity of the data and stores it in a database. It also cleans the stored data to remove unnecessary noise and privacy information. The cleaned data is converted into a training set and fed into an LLM (large-scale language model). This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The generated conversation model is stored in a repository so that it can be called up when needed.

[0135] The server also invokes the trained conversation model based on the user's request to generate a realistic response. The generated response is sent to the device and displayed to the user. Furthermore, if the user enters the next conversational phrase, the server analyzes it, generates the next response, and sends it to the device. This interaction takes place in real time, allowing the user to simulate a realistic conversation experience with a real person. Finally, these conversation logs are saved and used as subsequent training data.

[0136] Database

[0137] The database is a system that centrally stores and manages received conversation data and generated conversation models. The database stores received conversation data with an index, allowing for efficient management and search.

[0138] Specific examples

[0139] Let us say that a user wants to simulate an interview with a well-known professional.

[0140] 1. User: Logs in to the system using a terminal, selects a "famous professional" as the conversation partner, and selects the "interview format" as the situation.

[0141] 2. Server: Based on the user request, it invokes the "famous professional" conversation model and generates an initial response, such as "Hello, what would you like to talk about today?"

[0142] 3. Terminal: The initial response generated is displayed to the user.

[0143] 4. User: Enters the following question: "What event has had the greatest impact on your career?" and submits it to the server.

[0144] 5. Server: Based on the received question, the following response is generated: "The event that had the greatest impact on my career is..." and sent to the device. This is done in real time.

[0145] An example prompt might look like this:

[0146] "You are interviewing a well-known professional. The first question is, 'What has been the most influential event in your career so far?' Please continue with the following interview content."

[0147] Through this simulation, users can learn languages ​​through realistic conversational experiences, improving learning motivation and effectiveness.

[0148] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0149] Step 1:

[0150] Uploading conversation data

[0151] Terminal: The conversation data provider uses the terminal to select conversation data such as interview content or dialogue scenes, and clicks the upload button. The input is an audio file (e.g., .mp3) or a text file (e.g., .txt). The output is an upload request.

[0152] Step 2:

[0153] Receiving and storing data

[0154] Server: Receives conversation data sent from the device and temporarily stores it in storage. The input is the upload request from the device, and the output is the temporarily stored conversation data. Specifically, this includes checking the file size and format to ensure there are no inconsistencies.

[0155] Step 3:

[0156] Data Integrity Check

[0157] Server: Checks the integrity of the received data. The input is the temporarily saved conversation data, and the output is the status of whether it can be stored in the database. Specific operations include checking the completeness and consistency of the data.

[0158] Step 4:

[0159] Data storage and indexing

[0160] Server: Stores the integrity-checked data in a database and creates an index. The input is the checked conversation data, and the output is an indexed database entry, allowing for efficient searches.

[0161] Step 5:

[0162] Cleaning the data

[0163] Server: Cleans the conversation data stored in the database and removes unnecessary noise and privacy information. The input is the conversation data retrieved from the database, and the output is the clean data. Specifically, noise filtering is performed using a text mining library.

[0164] Step 6:

[0165] Creating a training set

[0166] Server: Converts the cleaned conversation data into a training set and feeds it to the LLM (large-scale language model). The input is clean data, and the output is the training set. Specifically, it normalizes and tokenizes the data.

[0167] Step 7:

[0168] Training a conversation model

[0169] Server: Trains the LLM using the training set. The input is the training set, and the output is the trained conversation model. Specific operations include updating the model parameters.

[0170] Step 8:

[0171] Saving the conversation model

[0172] Server: Stores the trained conversational model in a repository. The input is the trained conversational model, and the output is the model stored in the repository, so that it can be recalled and used later.

[0173] Step 9:

[0174] Accepting user requests

[0175] User: Logs in to the system through a terminal, selects a conversation partner and a situation, and requests a simulation. The input is the user's selection information, and the output is the request transmission.

[0176] Step 10:

[0177] Generate an initial response

[0178] Server: Based on the user's request, it calls the appropriate conversation model and generates an initial response. The input is the user's request, and the output is the generated initial response. The specific process is to input a prompt to the conversation model and get a response.

[0179] Step 11:

[0180] Viewing the response

[0181] Terminal: Displays the generated initial response on the screen. The input is the initial response from the server, and the output is the response displayed to the user.

[0182] Step 12:

[0183] Enter the next conversation phrase

[0184] User: Receives the initial response, enters the next conversational phrase, and sends it to the server through the terminal. The input is the next conversational phrase, and the output is a request to the server.

[0185] Step 13:

[0186] Producing the following response

[0187] Server: Analyzes the user's input, generates the next response, and sends it to the device. The input is the user's next conversation phrase, and the output is the generated response. This process is done in real time.

[0188] Step 14:

[0189] Save conversation logs

[0190] Server: Stores the interactions between the user and the conversation model as a log. The input is the data of each conversation turn, and the output is the saved conversation log. The log is also used as training data for subsequent tasks.

[0191] (Application example 1)

[0192] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0193] Today's language learners have limited opportunities to learn languages ​​efficiently through real-life conversation experiences with specific experts or celebrities. This issue can decrease learners' motivation and reduce learning effectiveness. Furthermore, content distribution services, in particular, lack platforms where users can enjoy interactive content. This leads to issues such as reduced user engagement and a decrease in the appeal of the service.

[0194] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0195] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) for training; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the terminal; a means for the server to store the interaction between the user and the conversation model as a log; a means for the server to store the generated conversation log and provide it for later review or additional learning; and a means for the terminal to display responses to the user based on the selected conversation situation. This allows users to learn languages ​​through realistic conversation experiences with celebrities and experts, improving learning motivation and effectiveness. It also enables content distribution services to provide users with more interactive and engaging content.

[0196] A "conversation data provider" is an individual or organization that is responsible for uploading conversation data from a terminal.

[0197] "Terminal" refers to a device used by a conversation data provider or user to access the system, including a smartphone, tablet, head-mounted display, computer, etc.

[0198] "Server" is a computer system that receives, stores, and analyzes conversation data and manages and serves trained models.

[0199] The "database" is a system that centrally stores and manages received conversation data and generated conversation models.

[0200] An "LLM (Large-Scale Language Model)" is an artificial intelligence model trained on a huge amount of text data and capable of generating natural-sounding conversations.

[0201] A "repository" is a data storage system where trained conversational models are stored.

[0202] A "user" is an individual or group that accesses the system, selects a conversation partner and situation, and requests a simulation.

[0203] A "conversational model" is an AI model that is generated based on a trained LLM and is used to generate responses corresponding to specific conversational scenarios.

[0204] A "realistic conversational experience" is a process that simulates natural interactions that are close to real conversations.

[0205] A "conversation log" is a record of the interactions between a user and a conversation model, and is used as subsequent training data.

[0206] "Review" is the process of reviewing past conversation logs to enhance learning effectiveness.

[0207] The present invention provides a system for language learning that allows users to engage in realistic conversational simulations with celebrities and professionals, and is particularly suited to devices such as smartphones, tablets, and head-mounted displays.

[0208] System Configuration

[0209] 1. Hardware Configuration

[0210] Terminal: A device on which users can enjoy interactive conversation simulations. This includes smartphones, tablets, and head-mounted displays.

[0211] Server: A computer system that stores and manages conversation data, trains LLMs (large-scale language models), and generates responses.

[0212] Database: A system that centrally stores and manages conversation data and generated conversation models.

[0213] 2. Software Configuration

[0214] Flask: A microframework for building web applications that handles server-side processing.

[0215] OpenAI API: Used to generate natural-sounding conversational responses based on user input using generative AI models.

[0216] Data processing and calculation

[0217] Receiving user input: The server receives phrases entered by the user on the device, including text input and voice input.

[0218] Prompt Generation: Based on the user's input, generate a prompt to continue the virtual conversation, for example, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0219] Response Generation: Uses OpenAI's ChatGPT API to generate appropriate responses based on the generated prompts.

[0220] Returning the response: The generated response is sent to the terminal and displayed to the user.

[0221] Saving conversation logs: Records of interactions between the user and the conversation model (conversation logs) are saved on the server.

[0222] Specific examples

[0223] If a user is trying to learn Japanese, the process would be as follows:

[0224] 1. User: I want to simulate a conversation with a famous actor.

[0225] 2. User Input: "Tell me about your passion for filmmaking."

[0226] 3. Server: Takes input and generates the following prompt:

[0227] Prompt: "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0228] 4. OpenAI API: Generate appropriate responses based on prompts.

[0229] 5. Terminal: Displays the generated response to the user.

[0230] In this way, users can effectively learn languages ​​through realistic conversation experiences with celebrities and experts.

[0231] The above is the "Mode for Carrying Out the Invention."

[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0233] Step 1:

[0234] The user selects a conversation partner and situation from the terminal and requests a simulation.

[0235] Input: User selection of conversation partner and situation.

[0236] Output: The request data to the server.

[0237] How it works: A user logs into the application using a smartphone, tablet, or head-mounted display, selects a conversation partner (e.g., a famous actor) and a situation for language learning, and then sends the selection to the server.

[0238] Step 2:

[0239] The server receives the request and invokes the conversation model.

[0240] Input: Request data from the user.

[0241] Output: The initial response of the conversation model.

[0242] Specific operation: Based on the received request, the server retrieves the corresponding conversation model from the database, for example, selects the conversation model of a famous actor, and generates an initial response.

[0243] Step 3:

[0244] The server receives the user's input and generates a prompt.

[0245] Input: A phrase entered by the user through the device (e.g., "Tell me about your passion for filmmaking.").

[0246] Output: The generated prompt statement.

[0247] Specific behavior: Based on the user's input phrase, the server generates a prompt sentence. The generated prompt sentence will be something like, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0248] Step 4:

[0249] The server calls the OpenAI API to generate a response based on the prompt.

[0250] Input: The generated prompt statement.

[0251] Output: The generated response.

[0252] Specific operation: The server uses the OpenAI API to generate an appropriate response based on the generated prompt. At this stage, the prompt is sent to the API and the response text is retrieved.

[0253] Step 5:

[0254] The server generates a response that is sent to the terminal and displayed to the user.

[0255] Input: The response obtained from the OpenAI API.

[0256] Output: The response text that is displayed on the terminal.

[0257] Specific operation: The server sends the generated response to the terminal, and the user confirms this response through the terminal. For example, the response displayed is "About your passion for filmmaking."

[0258] Step 6:

[0259] The server stores the interaction between the user and the conversation model as a log.

[0260] Input: The dialogue between the user and the conversation model.

[0261] Output: Saved conversation logs.

[0262] How it works: The server stores the interactions between the user and the conversation model in a database for subsequent learning and review. Conversation logs are automatically saved and made available for the user to refer to later.

[0263] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0264] The present invention combines an emotion engine with a language learning system that allows users to learn languages ​​through realistic conversation simulations, generating optimal responses according to the user's emotional state. This system aims to improve learners' motivation and learning effectiveness by providing simulations based on conversation data with celebrities and professionals, in particular.

[0265] System Configuration

[0266] The system includes the following major components:

[0267] 1. Terminal

[0268] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[0269] An interface for conversation data providers to upload conversation data.

[0270] An interface for users to request conversation simulations and view the simulation results.

[0271] 2. Server

[0272] The ability to receive, save, clean, and store conversation data in a database.

[0273] A function that trains LLMs (large-scale language models) based on saved conversation data.

[0274] Ability to save trained conversational models and recall them whenever needed.

[0275] The ability to use conversational models to generate realistic responses based on user requests.

[0276] Ability to recognize user emotions using an emotion engine and generate optimal responses.

[0277] A function that saves conversation logs and uses them as training data for future use.

[0278] 3. Database

[0279] A system that centrally stores and manages received conversation data and generated conversation models.

[0280] 4. Emotion Engine

[0281] An engine for recognizing emotions from user text input and voice data and optimizing responses based on that.

[0282] Collection and handling of conversation data

[0283] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[0284] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[0285] Training the model

[0286] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[0287] User Interaction

[0288] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[0289] Terminal: The selection is sent to the server as a request.

[0290] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[0291] Real-time conversation simulation and emotion recognition

[0292] Emotion engine: Recognizes emotions from the user's text input or voice data. The recognized emotion information is sent to the server.

[0293] Server: Optimizes responses based on the emotional information received from the emotion engine. Sends the generated responses to the device so that the conversation continues in real time.

[0294] User: Enter the following conversation phrase and send it to the server via the device. If an emotion is recognized, that information is also sent.

[0295] Server: Analyzes the user's input data and emotional information, and generates the following response, which is also sent to the device in an optimized form.

[0296] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[0297] Specific examples

[0298] A user simulates an interview with celebrity B.

[0299] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[0300] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[0301] 3. Terminal: The initial response is displayed to the user.

[0302] 4. User: Enter the following question and send it to the server via the terminal.

[0303] 5. Emotion engine: Recognizes emotions from the user's input data and also sends this information to the server.

[0304] 6. Server: Based on the emotion information received from the emotion engine, it generates and optimizes the next response and sends it to the terminal.

[0305] 7. Server: Saves the conversation log.

[0306] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[0307] The processing flow will be explained below.

[0308] Step 1: Upload your conversation data

[0309] Terminal: Conversation data providers use terminals to upload conversation data (voice or text) to the system.

[0310] Server: Receives conversation data and temporarily stores it in storage.

[0311] Step 2: Verify the integrity of the conversation data and store it

[0312] Server: Checks the integrity of the received conversation data and stores it in a database, adding appropriate metadata when storing it.

[0313] Step 3: Indexing

[0314] Server: Indexes the stored conversation data, enabling efficient search and management.

[0315] Step 4: Data Cleaning

[0316] Server: Cleans the stored conversation data to remove noise and privacy information.

[0317] Step 5: Generate a training set

[0318] Server: Converts the cleaned data into a training set and feeds it into the LLM (large-scale language model).

[0319] Step 6: Train and save the model

[0320] Server: Trains the LLM using the training set to learn the speaking style and phrases of a particular person.

[0321] Server: Stores the trained conversation model in a repository.

[0322] Step 7: Logging in users

[0323] User: Log in to the system and check your profile information.

[0324] Terminal: Sends login information to the server for authentication.

[0325] Step 8: Choose your conversation partner and situation

[0326] User: Selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the system.

[0327] Terminal: Sends the selection to the server.

[0328] Step 9: Receiving a request and generating an initial response

[0329] Server: Receives requests from users, invokes the selected conversation model, and generates the initial response.

[0330] Server: Sends the generated initial response to the terminal.

[0331] Step 10: View the initial response

[0332] Terminal: Displays the initial response received from the server to the user.

[0333] Step 11: Enter conversation phrases

[0334] User: Enter the following conversation phrase into your device:

[0335] Terminal: Sends the entered data to the server.

[0336] Step 12: Recognize emotions

[0337] Server: Uses an emotion engine to recognize emotions from user input data and voice data.

[0338] Step 13: Generate and optimize responses

[0339] Server: Generates the optimal response based on the user's input data and recognized emotional information.

[0340] Server: Sends the generated optimization response to the terminal.

[0341] Step 14: View the response

[0342] Terminal: displays the optimization response received from the server to the user.

[0343] Step 15: Save the conversation log

[0344] Server: Saves conversation logs with users in a database and uses them as training data for future use.

[0345] Specific examples

[0346] A user simulates an interview with celebrity B.

[0347] 1. Step 1:

[0348] Device: Celebrity B uploads his / her interview recording.

[0349] Server: Receives and temporarily stores conversation data.

[0350] 2. Step 2:

[0351] Server: Checks the integrity of the data and stores it in the database.

[0352] 3. Step 3:

[0353] Server: Creates indexes to simplify later searches.

[0354] 4. Steps 4-6:

[0355] Server: Performs data cleaning, generates training sets, and trains the LLM.

[0356] 5. Steps 7-8:

[0357] User: Log in to the system and select Celebrity B. Select the interview format.

[0358] Terminal: Sends the request contents to the server.

[0359] 6. Steps 9-10:

[0360] Server: Generates an initial response and sends it to the terminal.

[0361] Terminal: Display the initial response to the user.

[0362] 7. Steps 11-12:

[0363] User: Enter the following question and send it to the server via the terminal.

[0364] Server: Recognizes emotions using an emotion engine.

[0365] 8. Steps 13-14:

[0366] Server: Generates the optimal response based on the recognized emotional information and sends it to the device.

[0367] Terminal: Display the optimization response to the user.

[0368] 9. Step 15:

[0369] Server: Saves conversation logs and uses them as training data for future use.

[0370] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[0371] Example 2

[0372] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0373] Existing language learning systems have difficulty enabling users to effectively learn a language through realistic conversation simulations. Furthermore, they lack a mechanism for generating optimal responses based on the user's emotional state. As a result, learners' motivation and learning effectiveness may decrease.

[0374] The specification processing by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a conversation data provider to upload conversation data from an information terminal; a means for the server to save the received conversation data in a database; a means for the server to supply the saved conversation data to a large-scale language model and train it; a means for the server to save the trained conversation model in a repository; a means for a user to select a conversation partner and a situation from an information terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the information terminal; a means for an emotion engine to recognize emotions from the user's input data and send the emotion information to the server; a means for the server to optimize and generate a response based on the emotion information and send it to the information terminal; and a means for the server to save the interaction between the user and the conversation model as a log. This enables users to effectively learn language through realistic conversation simulations and improve their learning experience.

[0375] "Conversation Data Provider" means a person or organization that provides conversation data to the system.

[0376] An "information terminal" is an electronic device used to access a system and send and receive data. Examples include computers, smartphones, and tablets.

[0377] A "server" is a computer system responsible for receiving, processing, and storing conversation data.

[0378] "Conversation Data" refers to a record of a conversation provided as an audio file or text file.

[0379] A "database" is a storage system that allows the server to efficiently store, manage, and search conversation data received.

[0380] A "large-scale language model" is an algorithmic model for natural language processing that is trained on a large amount of text data.

[0381] "Training" refers to the process of feeding conversational data into a large-scale language model to learn specific speaking styles and phrases.

[0382] A "repository" is a storage system for storing trained conversational models.

[0383] "Profile information" refers to personal information and setting data that a user confirms after logging in to the system.

[0384] A "situation" is a specific situation or context in a conversation simulation, such as an interview format.

[0385] "Simulation" is the process of engaging in virtual interactions based on user-selected conversation partners and situations.

[0386] An "emotion engine" is an algorithm or system that recognizes emotions from user input data and provides the results to a server.

[0387] "Input data" refers to text and voice data provided by a user to a system.

[0388] "Emotional information" refers to data that indicates the emotional state of the user as recognized by the emotion engine.

[0389] "Log" refers to historical data that records interactions between users and conversation models.

[0390] "Optimization" refers to the process of appropriately adjusting the content and expression of a response based on recognized emotional information.

[0391] MODE FOR CARRYING OUT THE INVENTION

[0392] The present invention combines a system in which users learn languages ​​through realistic conversation simulations with an emotion engine to generate optimal responses according to the user's emotional state. The system includes the following main components:

[0393] 1. Terminal

[0394] A terminal is a device through which a data provider or user accesses the system. Examples include computers, smartphones, and tablets. Terminals are equipped with the following interfaces:

[0395] An interface for conversation data providers to upload conversation data

[0396] An interface for users to request conversational simulations and view the results of the simulations

[0397] 2. Server

[0398] The server receives, stores, cleans, and stores conversation data in a database. The server has the following functions:

[0399] 1. A function to receive conversation data and temporarily store it in storage

[0400] 2. A function to check the integrity of received data and store it in a database

[0401] 3. Ability to clean stored data and remove noise and private information

[0402] 4. The ability to convert cleaned data into a training set and feed it into a large-scale language model (LLM).

[0403] 5. Ability to save the conversation model generated through training and recall it whenever needed

[0404] 6. Ability to use conversational models to generate realistic responses based on user requests

[0405] 7. Ability to recognize user emotions using an emotion engine and generate optimal responses

[0406] 8. Ability to save conversation logs and use them as training data for future use

[0407] 3. Database

[0408] The database is a system that centrally stores and manages received conversation data and generated conversation models. Indexing enables efficient management and search.

[0409] 4. Emotion Engine

[0410] The emotion engine is an engine that recognizes emotions from user text input and voice data and optimizes responses based on those emotions.

[0411] Collection and processing of conversation data

[0412] Users use their devices to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. Audio files and text files can be used for this purpose. The device then sends the conversation data uploaded by the user to the server. The server checks the integrity of the received data and stores it in a database. Indexing allows for efficient management and searching.

[0413] Training the model

[0414] The server retrieves conversation data from the database and performs a data cleaning process. It converts the cleaned data into a training set and supplies it to the LLM. This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The server stores the generated conversation model in a repository and can be called up whenever needed.

[0415] Preparing for user interaction

[0416] A user logs in to the system using a terminal, checks their profile information, and selects a conversation partner and situation. The terminal then sends the selection to the server as a request. The server accepts the request, selects an appropriate conversation model from its saved storage, generates an initial response, and sends it to the terminal.

[0417] Real-time conversation simulation and emotion recognition

[0418] The user inputs the next question or dialogue phrase and sends it to the server via the device. The emotion engine analyzes the user's input data and recognizes emotions. The recognized emotion information is sent to the server. The server then optimizes the response based on the emotion information and sends the generated response to the device. This allows the user to effectively learn the language through realistic conversation simulation. The server saves these conversation logs and uses them as subsequent training data to improve the accuracy and quality of the system.

[0419] Specific examples

[0420] A user simulates an interview with a celebrity

[0421] 1. Terminal: The user logs in to the system and selects a celebrity to talk to. The situation is chosen to be an interview format.

[0422] 2. Server: Based on the user request, the celebrity conversation model is called and an initial response is generated, which reads, "Hello, celebrity. What would you like to talk about today?"

[0423] 3. Terminal: The initial response is displayed to the user. Type "Hello, celebrity. I'd like to ask you my next question. Tell me about your latest project."

[0424] 4. Emotion engine: Recognizes emotions from the user's input data and sends that information to the server. It is analyzed as "curious expression."

[0425] 5. Server: Based on the emotion information received from the emotion engine, the server generates and optimizes a response and sends it to the device. The response generated is "Thank you. Now, I'll tell you about my latest project..." and is displayed to the user.

[0426] 6. Server: Stores user interaction logs and uses them as training data for subsequent use to improve the accuracy of the system.

[0427] This allows users to advance their language learning while enjoying a realistic interview experience.

[0428] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0429] Specific processing steps of the system

[0430] Step 1:

[0431] A user logs into a system using an information terminal. The input is the user's authentication information (e.g., username, password), and the output is the start of a session for the authenticated user. To access the system, the user enters the appropriate authentication information. The system receives this, authenticates, and then starts the user session.

[0432] Step 2:

[0433] Users upload their own conversation data from their information terminals. The input is an audio file or a text file, and the output is the conversation data sent to the server. Using the system's interface, conversation data providers select a data file and upload it to the server.

[0434] Step 3:

[0435] The server temporarily stores the received conversation data in storage. The input is the conversation data sent from the terminal, and the output is the data temporarily stored in storage. After receiving the data, the server verifies it, performs an integrity check, and then stores the data in storage.

[0436] Step 4:

[0437] The server stores the saved conversation data in a database. The input is the data stored in temporary storage, and the output is the data indexed in the database. When storing the data in the database, an index is created to enable efficient management and search.

[0438] Step 5:

[0439] The server retrieves the conversation data from the database and runs a data cleaning process. The input is the raw data in the database, and the output is the cleaned data. The dataset is cleaned using algorithms to remove noise and privacy information.

[0440] Step 6:

[0441] The server converts the cleaned conversation data into a training set and feeds it to an LLM (large-scale language model). The input is the cleaned data, and the output is the training set. This completes the dataset for training the conversation model.

[0442] Step 7:

[0443] The server saves the conversation model generated by training in a repository. The input is the conversation model generated by training, and the output is the model saved in the repository. The generated model is saved in storage so that it can be used in response to any request.

[0444] Step 8:

[0445] The user selects a conversation partner and a situation using an information terminal. The input is the selection information of the conversation partner and the situation, and the output is request data to the server. The details of the conversation simulation are specified through the interface, and the information is sent to the system.

[0446] Step 9:

[0447] The server calls the appropriate stored conversation model based on the user's request and generates an initial response. The input is the user's request data, and the output is the generated initial response. The specified model is loaded and an initial response according to the request is generated.

[0448] Step 10:

[0449] The server sends the generated initial response to the information terminal. The input is the generated initial response, and the output is the response displayed on the terminal. The generated response is forwarded to the user for display.

[0450] Step 11:

[0451] The user inputs the next question or dialogue phrase and sends it to the server through the terminal. The input is the user's next question or dialogue phrase, and the output is the data sent to the server. The user inputs the next step to continue the dialogue.

[0452] Step 12:

[0453] The emotion engine analyzes the user's input data and recognizes emotions. The input is the user's next question or dialogue phrase, and the output is the recognized emotion information. It analyzes the provided data and recognizes the user's emotional state.

[0454] Step 13:

[0455] The emotion engine sends the recognized emotion information to the server. The input is the recognized emotion information, and the output is the data sent to the server. The analysis results are transferred to the server.

[0456] Step 14:

[0457] The server optimizes and generates a response based on the emotional information. The input is the emotional information and the user's next question, and the output is an optimized response. The response content is adjusted and generated according to the recognized emotion.

[0458] Step 15:

[0459] The server sends the generated response to the terminal and displays it to the user. The input is the generated response and the output is the response displayed on the terminal. The server forwards the response to the user and continues the dialogue.

[0460] Step 16:

[0461] The server stores the interaction between the user and the conversational model as a log. The input is the interaction data between the user and the conversational model, and the output is the stored log. The stored log is used for subsequent training.

[0462] (Application example 2)

[0463] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0464] Conventional language learning systems have had difficulty providing optimal responses based on conversational realism and emotions. It has also been difficult for users to engage in natural conversations in virtual stores while selecting products and receiving information in real time. This has led to problems such as reduced learning effectiveness and user satisfaction.

[0465] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0466] In this invention, the server includes means for analyzing the emotional state of the user using an emotion engine and optimizing a response based on that information, means for providing the optimized response to the user's terminal in real time, and means for the user to select products and provide information in the virtual environment using the terminal, thereby enabling the user to obtain product information while engaging in natural conversation according to their emotions in real time in the virtual store.

[0467] A "conversation data provider" is a person or organization that uploads conversation data from a terminal.

[0468] "Terminal" means a device (e.g., computer, smartphone, tablet) used by a conversation data provider or user to access the system.

[0469] The "server" is the core computer system that receives and stores conversation data, serves it as training data, and ultimately generates responses.

[0470] A "database" is a system for storing and managing received conversation data and training data.

[0471] An "LLM (Large Scale Language Model)" is a language model trained using a large number of datasets to achieve natural language processing.

[0472] "Training" is the process of optimizing a large-scale language model using stored conversational data.

[0473] A "repository" is a system for storing and managing trained conversation models.

[0474] "User" means a person or entity that uses a terminal to access the system and use the conversation simulation.

[0475] An "emotion engine" is a technology that analyzes a user's emotional state and optimizes responses based on that information.

[0476] A "virtual environment" is a simulated environment that allows users to have a realistic experience within a digital space.

[0477] "Product selection" is the action of a user selecting a product of interest within a virtual environment.

[0478] "Information provision" means providing detailed information about the product selected by the user through a virtual assistant or the like.

[0479] "Real-time" means that processing is done immediately, without delay, and the results are provided to the user.

[0480] The following describes in detail an embodiment of the invention in a virtual store customer service system. First, the process of generating a program for the system that realizes this application example and its specific operation will be described.

[0481] Overall system configuration

[0482] The system consists of the following main elements:

[0483] 1. Terminal: A device through which a user accesses the system, such as a smartphone or head-mounted display.

[0484] 2. Server: The core computer system that receives, stores, trains, and generates responses from conversation data.

[0485] 3. Database: A system that stores and manages conversation data and conversation models.

[0486] 4. Emotion Engine: Technology that analyzes the user's emotional state and optimizes responses based on that information.

[0487] Detailed Operation

[0488] 1. User visits the virtual store:

[0489] By launching the application using a device (smartphone or head-mounted display), the user enters the 3D environment of the virtual store.

[0490] 2. Product Selection and Information:

[0491] Once the user selects the product of interest, product information is provided by the virtual assistant (trained conversation model).

[0492] 3. Real-time conversation simulation:

[0493] When a user asks a question to a virtual assistant, the emotion engine analyzes the question and the user's emotional state. The emotion engine's API (e.g., IBM Watson Tone Analyzer, Microsoft Azure Emotion API) analyzes the user's voice data and text to recognize the emotional state.

[0494] 4. Optimal response generation:

[0495] Based on the recognized emotional state, a large-scale language model (e.g., OpenAI's GPT series) generates an optimal response. The server sends a request to the triggered model and obtains the inference result.

[0496] 5. Response display:

[0497] The generated responses are displayed to the user through the device, and the conversation logs are saved by the server and used as training data for future use.

[0498] Specific hardware and software

[0499] Hardware

[0500] Smartphones: iPhone, Android, etc.

[0501] Head-mounted displays: Oculus Quest 2, HoloLens 2, etc.

[0502] software

[0503] Front-end frameworks: Unity, React Native, etc.

[0504] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.

[0505] Large-scale language models: OpenAI GPT-3, GPT-4, etc.

[0506] Database: MongoDB, MySQL, etc.

[0507] Server-side frameworks: Node.js, Django, etc.

[0508] Examples of concrete examples and prompts

[0509] Here's how the system works:

[0510] The user puts on a head-mounted display, enters a virtual store, and approaches specific products.

[0511] Example prompt: "What features does this new smartwatch have?"

[0512] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[0513] This allows users to obtain product information while engaging in natural, emotionally-responsive dialogue in real time in a virtual store.

[0514] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0515] Step 1:

[0516] Access to the system and entry to the virtual store

[0517] The user launches the application using a smartphone or head-mounted display, which accesses the 3D environment of the virtual store, where the user logs in to the system and the initial interface is displayed.

[0518] Input: User login information, device type

[0519] Data processing / calculation: User authentication, display of virtual store using 3D rendering engine

[0520] Output: Login success message, virtual store UI display

[0521] Step 2:

[0522] Product selection and initial information provision

[0523] The user moves through the virtual environment and selects a specific product. The device sends this selection information to the server, which retrieves the product's initial information from a database and displays it.

[0524] Input: ID of the product selected by the user

[0525] Data processing / calculation: database query for product information, information format conversion

[0526] Output: Product details (text, images, etc.)

[0527] Step 3:

[0528] Start the conversation simulation

[0529] The user inputs a question related to the product, and the device sends the question to the server, which then uses an emotion engine to analyze the user's emotional state.

[0530] Input: User text input (question content)

[0531] Data processing / calculation: Emotion analysis using emotion engines (e.g., IBM Watson Tone Analyzer)

[0532] Output: Emotional state data (e.g., surprise, joy, dissatisfaction)

[0533] Step 4:

[0534] Generating the best response

[0535] The server uses the emotional state data and the user's question to trigger a large-scale language model (e.g., OpenAI GPT-4) to generate an optimal response.

[0536] Input: User question, emotional state data

[0537] Data processing / calculation: Prompt generation and response generation for large-scale language models

[0538] Output: Optimized response text

[0539] Step 5:

[0540] Viewing the response

[0541] The generated response text is sent from the server to the terminal, which displays the response to the user.

[0542] Input: Response text data from the server

[0543] Data processing / calculation: rendering text data, updating the interface

[0544] Output: User confirms response text

[0545] Step 6:

[0546] Save conversation logs

[0547] The conversational exchanges are logged and stored by the server and used as training data for subsequent use.

[0548] Input: Conversation data between the user and the virtual assistant

[0549] Data processing / calculation: Logging to database, indexing

[0550] Output: Saved conversation log

[0551] Examples:

[0552] Prompt Sentence Examples

[0553] A user asks: "What features does this new smartwatch have?"

[0554] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[0555] This step allows users to receive appropriate responses in real time based on their emotions, enabling an efficient and satisfying virtual store experience.

[0556] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0557] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0558] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0559] [Second embodiment]

[0560] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0561] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0562] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0563] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0564] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0566] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0567] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0568] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0569] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0570] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0571] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0572] The present invention is a system that allows users to learn a language through realistic conversation simulations. This system provides simulations based on conversation data with famous people and professionals, and aims to improve learners' motivation and learning effectiveness.

[0573] System Configuration

[0574] The system includes the following major components:

[0575] 1. Terminal

[0576] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[0577] An interface for conversation data providers to upload conversation data.

[0578] An interface for users to request conversation simulations and view the simulation results.

[0579] 2. Server

[0580] The ability to receive, save, clean, and store conversation data in a database.

[0581] A function that trains LLMs (large-scale language models) based on saved conversation data.

[0582] Ability to save trained conversational models and recall them whenever needed.

[0583] The ability to use conversational models to generate realistic responses based on user requests.

[0584] A function that saves conversation logs and uses them as training data for future use.

[0585] 3. Database

[0586] A system that centrally stores and manages received conversation data and generated conversation models.

[0587] Collection and handling of conversation data

[0588] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[0589] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[0590] Training the model

[0591] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[0592] User Interaction

[0593] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[0594] Terminal: The selection is sent to the server as a request.

[0595] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[0596] Real-time conversation simulation

[0597] User: Enter the following conversation phrase and send it to the server via the device.

[0598] Server: Analyzes the user's input data, generates the next response, and sends it to the device. This interaction occurs in real time, allowing the user to simulate a realistic conversation experience with a real person.

[0599] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[0600] Specific examples

[0601] A user simulates an interview with celebrity B.

[0602] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[0603] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[0604] 3. Terminal: The initial response is displayed to the user.

[0605] 4. User: Enter the following question and submit it to the server:

[0606] 5. Server: Generates the following response and sends it to the device:

[0607] 6. Server: Stores the conversation log.

[0608] In this way, users can advance their language learning while simulating a real-life interview experience with celebrity B.

[0609] The processing flow will be explained below.

[0610] Step 1: Upload your conversation data

[0611] Terminal: Conversation data providers use terminals to upload conversation data (audio files or text files) to the system.

[0612] Server: Stores the received conversation data in temporary storage and checks the data integrity.

[0613] Step 2: Storing and indexing conversation data

[0614] Server: Stores the verified conversation data in a database.

[0615] Server: Creates an index of the stored data to facilitate future search and management.

[0616] Step 3: Cleaning the conversation data

[0617] Server: Cleans the stored conversation data, removing noise and privacy information.

[0618] Step 4: Generate a training set

[0619] Server: Converts the cleaned data into a training set for LLMs (large-scale language models).

[0620] Step 5: Train the model

[0621] Server: Feeds the training set to the LLM, allowing it to learn specific person speaking styles and phrases.

[0622] Server: Stores the trained model in a repository.

[0623] Step 6: User login and profile verification

[0624] User: Log in to the system and check your profile information.

[0625] Terminal: Sends the user's login information to the server for authentication.

[0626] Step 7: Choose your conversation partner and situation

[0627] User: Select a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the terminal.

[0628] Terminal: Sends the selection to the server as a request.

[0629] Step 8: Generate an initial response

[0630] Server: Based on the user's request, it invokes the appropriate conversation model to generate the initial response.

[0631] Server: Sends the generated initial response to the terminal.

[0632] Step 9: View the initial response

[0633] Terminal: Displays to the user the initial response received from the server.

[0634] Step 10: Enter conversation phrases

[0635] User: Enter the following conversation phrase into your device:

[0636] Terminal: Sends user input data to the server.

[0637] Step 11: Generate and Send Response

[0638] Server: Parses the user's input data and generates the following response:

[0639] Server: Sends the generated response to the terminal.

[0640] Step 12: Save the conversation log

[0641] Server: Conversation logs with the user are stored in a database and used as training data for future use.

[0642] Example 1

[0643] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0644] Language learning is becoming increasingly important in modern times, and learning through real conversations is particularly effective. However, opportunities to experience real conversations are limited, making it difficult to learn through conversation simulations with specific celebrities or professionals. Furthermore, systems that generate responses in real time in response to user requests are not common, making it difficult to provide a conversation experience that is close to reality. This has led to problems such as insufficient improvement in user motivation and effectiveness of learning.

[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0646] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) and train it; a means for the server to store the trained conversation model in a repository; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response and send it to the terminal; a real-time conversation simulation means for the user to input the next conversation phrase and the server to generate a response and send it to the terminal; and a means for the server to store the interaction between the user and the conversation model as a log. This enables language learning through realistic conversation simulations with specific celebrities or professionals, thereby improving the motivation and effectiveness of users' learning.

[0647] "Conversation data provider" refers to a user or individual who provides conversation data to the conversation simulation system.

[0648] "Terminal" refers to a device, such as a computer, smartphone, or tablet, used by a conversation data provider or user to access the system.

[0649] "Conversation data" refers to materials such as audio files and text files used for conversation simulation.

[0650] "Server" refers to a computer system that receives, stores, processes data, trains models, etc.

[0651] "Database" refers to a system for efficiently storing, managing, and searching conversation data.

[0652] "LLM (Large-scale Language Model)" refers to an AI model for natural language processing that is trained based on massive amounts of conversational data.

[0653] "Training Set" refers to the collection of cleaned conversational data used to train the LLM.

[0654] "Repository" refers to a storage system for storing trained conversational models.

[0655] "User" refers to an individual who utilizes the conversation simulation system to simulate a real conversation experience.

[0656] "Situation" refers to the conversation format or setting selected by the user, such as an interview format.

[0657] A "request" refers to a request or instruction a user makes to a system.

[0658] "Response" means any reply or response content generated by LLM to a User's request.

[0659] "Real-time conversation simulation" refers to a process in which the server instantly generates and returns a response to a conversation phrase entered by the user.

[0660] "Conversation log" refers to data that records the interactions between a user and a conversation model.

[0661] "Cleaning" refers to the process of removing unnecessary noise and private information from conversation data.

[0662] The present invention is a system that allows users to learn a language through realistic conversation simulations. One feature of this system is that it provides simulations based on conversation data with famous people and professionals, thereby improving learners' motivation and learning effectiveness. An embodiment of the present invention is described in detail below.

[0663] This system mainly consists of three components: a terminal, a server, and a database.

[0664] Terminal

[0665] Terminals are devices through which conversation data providers and users access the system, and examples include computers, smartphones, and tablets. Conversation data providers use these terminals to access an interface for uploading conversation data such as audio files and text files to the system. Users can also request conversation simulations and view the simulation results through the terminals.

[0666] server

[0667] The server is responsible for the core processing of the system. First, it receives conversation data uploaded from the device and temporarily stores it in storage. Next, it checks the integrity of the data and stores it in a database. It also cleans the stored data to remove unnecessary noise and privacy information. The cleaned data is converted into a training set and fed into an LLM (large-scale language model). This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The generated conversation model is stored in a repository so that it can be called up when needed.

[0668] The server also invokes the trained conversation model based on the user's request to generate a realistic response. The generated response is sent to the device and displayed to the user. Furthermore, if the user enters the next conversational phrase, the server analyzes it, generates the next response, and sends it to the device. This interaction takes place in real time, allowing the user to simulate a realistic conversation experience with a real person. Finally, these conversation logs are saved and used as subsequent training data.

[0669] Database

[0670] The database is a system that centrally stores and manages received conversation data and generated conversation models. The database stores received conversation data with an index, allowing for efficient management and search.

[0671] Specific examples

[0672] Let us say that a user wants to simulate an interview with a well-known professional.

[0673] 1. User: Logs in to the system using a terminal, selects a "famous professional" as the conversation partner, and selects the "interview format" as the situation.

[0674] 2. Server: Based on the user request, it invokes the "famous professional" conversation model and generates an initial response, such as "Hello, what would you like to talk about today?"

[0675] 3. Terminal: The initial response generated is displayed to the user.

[0676] 4. User: Enters the following question: "What event has had the greatest impact on your career?" and submits it to the server.

[0677] 5. Server: Based on the received question, the following response is generated: "The event that had the greatest impact on my career is..." and sent to the device. This is done in real time.

[0678] An example prompt might look like this:

[0679] "You are interviewing a well-known professional. The first question is, 'What has been the most influential event in your career so far?' Please continue with the following interview content."

[0680] Through this simulation, users can learn languages ​​through realistic conversational experiences, improving learning motivation and effectiveness.

[0681] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0682] Step 1:

[0683] Uploading conversation data

[0684] Terminal: The conversation data provider uses the terminal to select conversation data such as interview content or dialogue scenes, and clicks the upload button. The input is an audio file (e.g., .mp3) or a text file (e.g., .txt). The output is an upload request.

[0685] Step 2:

[0686] Receiving and storing data

[0687] Server: Receives conversation data sent from the device and temporarily stores it in storage. The input is the upload request from the device, and the output is the temporarily stored conversation data. Specifically, this includes checking the file size and format to ensure there are no inconsistencies.

[0688] Step 3:

[0689] Data Integrity Check

[0690] Server: Checks the integrity of the received data. The input is the temporarily saved conversation data, and the output is the status of whether it can be stored in the database. Specific operations include checking the completeness and consistency of the data.

[0691] Step 4:

[0692] Data storage and indexing

[0693] Server: Stores the integrity-checked data in a database and creates an index. The input is the checked conversation data, and the output is an indexed database entry, allowing for efficient searches.

[0694] Step 5:

[0695] Cleaning the data

[0696] Server: Cleans the conversation data stored in the database and removes unnecessary noise and privacy information. The input is the conversation data retrieved from the database, and the output is the clean data. Specifically, noise filtering is performed using a text mining library.

[0697] Step 6:

[0698] Creating a training set

[0699] Server: Converts the cleaned conversation data into a training set and feeds it to the LLM (large-scale language model). The input is clean data, and the output is the training set. Specifically, it normalizes and tokenizes the data.

[0700] Step 7:

[0701] Training a conversation model

[0702] Server: Trains the LLM using the training set. The input is the training set, and the output is the trained conversation model. Specific operations include updating the model parameters.

[0703] Step 8:

[0704] Saving the conversation model

[0705] Server: Stores the trained conversational model in a repository. The input is the trained conversational model, and the output is the model stored in the repository, so that it can be recalled and used later.

[0706] Step 9:

[0707] Accepting user requests

[0708] User: Logs in to the system through a terminal, selects a conversation partner and a situation, and requests a simulation. The input is the user's selection information, and the output is the request transmission.

[0709] Step 10:

[0710] Generate an initial response

[0711] Server: Based on the user's request, it calls the appropriate conversation model and generates an initial response. The input is the user's request, and the output is the generated initial response. The specific process is to input a prompt to the conversation model and get a response.

[0712] Step 11:

[0713] Viewing the response

[0714] Terminal: Displays the generated initial response on the screen. The input is the initial response from the server, and the output is the response displayed to the user.

[0715] Step 12:

[0716] Enter the next conversation phrase

[0717] User: Receives the initial response, enters the next conversational phrase, and sends it to the server through the terminal. The input is the next conversational phrase, and the output is a request to the server.

[0718] Step 13:

[0719] Producing the following response

[0720] Server: Analyzes the user's input, generates the next response, and sends it to the device. The input is the user's next conversation phrase, and the output is the generated response. This process is done in real time.

[0721] Step 14:

[0722] Save conversation logs

[0723] Server: Stores the interactions between the user and the conversation model as a log. The input is the data of each conversation turn, and the output is the saved conversation log. The log is also used as training data for subsequent tasks.

[0724] (Application example 1)

[0725] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0726] Today's language learners have limited opportunities to learn languages ​​efficiently through real-life conversation experiences with specific experts or celebrities. This issue can decrease learners' motivation and reduce learning effectiveness. Furthermore, content distribution services, in particular, lack platforms where users can enjoy interactive content. This leads to issues such as reduced user engagement and a decrease in the appeal of the service.

[0727] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0728] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) for training; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the terminal; a means for the server to store the interaction between the user and the conversation model as a log; a means for the server to store the generated conversation log and provide it for later review or additional learning; and a means for the terminal to display responses to the user based on the selected conversation situation. This allows users to learn languages ​​through realistic conversation experiences with celebrities and experts, improving learning motivation and effectiveness. It also enables content distribution services to provide users with more interactive and engaging content.

[0729] A "conversation data provider" is an individual or organization that is responsible for uploading conversation data from a terminal.

[0730] "Terminal" refers to a device used by a conversation data provider or user to access the system, including a smartphone, tablet, head-mounted display, computer, etc.

[0731] "Server" is a computer system that receives, stores, and analyzes conversation data and manages and serves trained models.

[0732] The "database" is a system that centrally stores and manages received conversation data and generated conversation models.

[0733] An "LLM (Large-Scale Language Model)" is an artificial intelligence model trained on a huge amount of text data and capable of generating natural-sounding conversations.

[0734] A "repository" is a data storage system where trained conversational models are stored.

[0735] A "user" is an individual or group that accesses the system, selects a conversation partner and situation, and requests a simulation.

[0736] A "conversational model" is an AI model that is generated based on a trained LLM and is used to generate responses corresponding to specific conversational scenarios.

[0737] A "realistic conversational experience" is a process that simulates natural interactions that are close to real conversations.

[0738] A "conversation log" is a record of the interactions between a user and a conversation model, and is used as subsequent training data.

[0739] "Review" is the process of reviewing past conversation logs to enhance learning effectiveness.

[0740] The present invention provides a system for language learning that allows users to engage in realistic conversational simulations with celebrities and professionals, and is particularly suited to devices such as smartphones, tablets, and head-mounted displays.

[0741] System Configuration

[0742] 1. Hardware Configuration

[0743] Terminal: A device on which users can enjoy interactive conversation simulations. This includes smartphones, tablets, and head-mounted displays.

[0744] Server: A computer system that stores and manages conversation data, trains LLMs (large-scale language models), and generates responses.

[0745] Database: A system that centrally stores and manages conversation data and generated conversation models.

[0746] 2. Software Configuration

[0747] Flask: A microframework for building web applications that handles server-side processing.

[0748] OpenAI API: Used to generate natural-sounding conversational responses based on user input using generative AI models.

[0749] Data processing and calculation

[0750] Receiving user input: The server receives phrases entered by the user on the device, including text input and voice input.

[0751] Prompt Generation: Based on the user's input, generate a prompt to continue the virtual conversation, for example, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0752] Response Generation: Uses OpenAI's ChatGPT API to generate appropriate responses based on the generated prompts.

[0753] Returning the response: The generated response is sent to the terminal and displayed to the user.

[0754] Saving conversation logs: Records of interactions between the user and the conversation model (conversation logs) are saved on the server.

[0755] Specific examples

[0756] If a user is trying to learn Japanese, the process would be as follows:

[0757] 1. User: I want to simulate a conversation with a famous actor.

[0758] 2. User Input: "Tell me about your passion for filmmaking."

[0759] 3. Server: Takes input and generates the following prompt:

[0760] Prompt: "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0761] 4. OpenAI API: Generate appropriate responses based on prompts.

[0762] 5. Terminal: Displays the generated response to the user.

[0763] In this way, users can effectively learn languages ​​through realistic conversation experiences with celebrities and experts.

[0764] The above is the "Mode for Carrying Out the Invention."

[0765] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0766] Step 1:

[0767] The user selects a conversation partner and situation from the terminal and requests a simulation.

[0768] Input: User selection of conversation partner and situation.

[0769] Output: The request data to the server.

[0770] How it works: A user logs into the application using a smartphone, tablet, or head-mounted display, selects a conversation partner (e.g., a famous actor) and a situation for language learning, and then sends the selection to the server.

[0771] Step 2:

[0772] The server receives the request and invokes the conversation model.

[0773] Input: Request data from the user.

[0774] Output: The initial response of the conversation model.

[0775] Specific operation: Based on the received request, the server retrieves the corresponding conversation model from the database, for example, selects the conversation model of a famous actor, and generates an initial response.

[0776] Step 3:

[0777] The server receives the user's input and generates a prompt.

[0778] Input: A phrase entered by the user through the device (e.g., "Tell me about your passion for filmmaking.").

[0779] Output: The generated prompt statement.

[0780] Specific behavior: Based on the user's input phrase, the server generates a prompt sentence. The generated prompt sentence will be something like, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[0781] Step 4:

[0782] The server calls the OpenAI API to generate a response based on the prompt.

[0783] Input: The generated prompt statement.

[0784] Output: The generated response.

[0785] Specific operation: The server uses the OpenAI API to generate an appropriate response based on the generated prompt. At this stage, the prompt is sent to the API and the response text is retrieved.

[0786] Step 5:

[0787] The server generates a response that is sent to the terminal and displayed to the user.

[0788] Input: The response obtained from the OpenAI API.

[0789] Output: The response text that is displayed on the terminal.

[0790] Specific operation: The server sends the generated response to the terminal, and the user confirms this response through the terminal. For example, the response displayed is "About your passion for filmmaking."

[0791] Step 6:

[0792] The server stores the interaction between the user and the conversation model as a log.

[0793] Input: The dialogue between the user and the conversation model.

[0794] Output: Saved conversation logs.

[0795] How it works: The server stores the interactions between the user and the conversation model in a database for subsequent learning and review. Conversation logs are automatically saved and made available for the user to refer to later.

[0796] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0797] The present invention combines an emotion engine with a language learning system that allows users to learn languages ​​through realistic conversation simulations, generating optimal responses according to the user's emotional state. This system aims to improve learners' motivation and learning effectiveness by providing simulations based on conversation data with celebrities and professionals, in particular.

[0798] System Configuration

[0799] The system includes the following major components:

[0800] 1. Terminal

[0801] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[0802] An interface for conversation data providers to upload conversation data.

[0803] An interface for users to request conversation simulations and view the simulation results.

[0804] 2. Server

[0805] The ability to receive, save, clean, and store conversation data in a database.

[0806] A function that trains LLMs (large-scale language models) based on saved conversation data.

[0807] Ability to save trained conversational models and recall them whenever needed.

[0808] The ability to use conversational models to generate realistic responses based on user requests.

[0809] Ability to recognize user emotions using an emotion engine and generate optimal responses.

[0810] A function that saves conversation logs and uses them as training data for future use.

[0811] 3. Database

[0812] A system that centrally stores and manages received conversation data and generated conversation models.

[0813] 4. Emotion Engine

[0814] An engine for recognizing emotions from user text input and voice data and optimizing responses based on that.

[0815] Collection and handling of conversation data

[0816] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[0817] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[0818] Training the model

[0819] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[0820] User Interaction

[0821] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[0822] Terminal: The selection is sent to the server as a request.

[0823] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[0824] Real-time conversation simulation and emotion recognition

[0825] Emotion engine: Recognizes emotions from the user's text input or voice data. The recognized emotion information is sent to the server.

[0826] Server: Optimizes responses based on the emotional information received from the emotion engine. Sends the generated responses to the device so that the conversation continues in real time.

[0827] User: Enter the following conversation phrase and send it to the server via the device. If an emotion is recognized, that information is also sent.

[0828] Server: Analyzes the user's input data and emotional information, and generates the following response, which is also sent to the device in an optimized form.

[0829] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[0830] Specific examples

[0831] A user simulates an interview with celebrity B.

[0832] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[0833] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[0834] 3. Terminal: The initial response is displayed to the user.

[0835] 4. User: Enter the following question and send it to the server via the terminal.

[0836] 5. Emotion engine: Recognizes emotions from the user's input data and also sends this information to the server.

[0837] 6. Server: Based on the emotion information received from the emotion engine, it generates and optimizes the next response and sends it to the terminal.

[0838] 7. Server: Saves the conversation log.

[0839] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[0840] The processing flow will be explained below.

[0841] Step 1: Upload your conversation data

[0842] Terminal: Conversation data providers use terminals to upload conversation data (voice or text) to the system.

[0843] Server: Receives conversation data and temporarily stores it in storage.

[0844] Step 2: Verify the integrity of the conversation data and store it

[0845] Server: Checks the integrity of the received conversation data and stores it in a database, adding appropriate metadata when storing it.

[0846] Step 3: Indexing

[0847] Server: Indexes the stored conversation data, enabling efficient search and management.

[0848] Step 4: Data Cleaning

[0849] Server: Cleans the stored conversation data to remove noise and privacy information.

[0850] Step 5: Generate a training set

[0851] Server: Converts the cleaned data into a training set and feeds it into the LLM (large-scale language model).

[0852] Step 6: Train and save the model

[0853] Server: Trains the LLM using the training set to learn the speaking style and phrases of a particular person.

[0854] Server: Stores the trained conversation model in a repository.

[0855] Step 7: Logging in users

[0856] User: Log in to the system and check your profile information.

[0857] Terminal: Sends login information to the server for authentication.

[0858] Step 8: Choose your conversation partner and situation

[0859] User: Selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the system.

[0860] Terminal: Sends the selection to the server.

[0861] Step 9: Receiving a request and generating an initial response

[0862] Server: Receives requests from users, invokes the selected conversation model, and generates the initial response.

[0863] Server: Sends the generated initial response to the terminal.

[0864] Step 10: View the initial response

[0865] Terminal: Displays the initial response received from the server to the user.

[0866] Step 11: Enter conversation phrases

[0867] User: Enter the following conversation phrase into your device:

[0868] Terminal: Sends the entered data to the server.

[0869] Step 12: Recognize emotions

[0870] Server: Uses an emotion engine to recognize emotions from user input data and voice data.

[0871] Step 13: Generate and optimize responses

[0872] Server: Generates the optimal response based on the user's input data and recognized emotional information.

[0873] Server: Sends the generated optimization response to the terminal.

[0874] Step 14: View the response

[0875] Terminal: displays the optimization response received from the server to the user.

[0876] Step 15: Save the conversation log

[0877] Server: Saves conversation logs with users in a database and uses them as training data for future use.

[0878] Specific examples

[0879] A user simulates an interview with celebrity B.

[0880] 1. Step 1:

[0881] Device: Celebrity B uploads his / her interview recording.

[0882] Server: Receives and temporarily stores conversation data.

[0883] 2. Step 2:

[0884] Server: Checks the integrity of the data and stores it in the database.

[0885] 3. Step 3:

[0886] Server: Creates indexes to simplify later searches.

[0887] 4. Steps 4-6:

[0888] Server: Performs data cleaning, generates training sets, and trains the LLM.

[0889] 5. Steps 7-8:

[0890] User: Log in to the system and select Celebrity B. Select the interview format.

[0891] Terminal: Sends the request contents to the server.

[0892] 6. Steps 9-10:

[0893] Server: Generates an initial response and sends it to the terminal.

[0894] Terminal: Display the initial response to the user.

[0895] 7. Steps 11-12:

[0896] User: Enter the following question and send it to the server via the terminal.

[0897] Server: Recognizes emotions using an emotion engine.

[0898] 8. Steps 13-14:

[0899] Server: Generates the optimal response based on the recognized emotional information and sends it to the device.

[0900] Terminal: Display the optimization response to the user.

[0901] 9. Step 15:

[0902] Server: Saves conversation logs and uses them as training data for future use.

[0903] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[0904] Example 2

[0905] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0906] Existing language learning systems have difficulty enabling users to effectively learn a language through realistic conversation simulations. Furthermore, they lack a mechanism for generating optimal responses based on the user's emotional state. As a result, learners' motivation and learning effectiveness may decrease.

[0907] The specification processing by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a conversation data provider to upload conversation data from an information terminal; a means for the server to save the received conversation data in a database; a means for the server to supply the saved conversation data to a large-scale language model and train it; a means for the server to save the trained conversation model in a repository; a means for a user to select a conversation partner and a situation from an information terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the information terminal; a means for an emotion engine to recognize emotions from the user's input data and send the emotion information to the server; a means for the server to optimize and generate a response based on the emotion information and send it to the information terminal; and a means for the server to save the interaction between the user and the conversation model as a log. This enables users to effectively learn language through realistic conversation simulations and improve their learning experience.

[0908] "Conversation Data Provider" means a person or organization that provides conversation data to the system.

[0909] An "information terminal" is an electronic device used to access a system and send and receive data. Examples include computers, smartphones, and tablets.

[0910] A "server" is a computer system responsible for receiving, processing, and storing conversation data.

[0911] "Conversation Data" refers to a record of a conversation provided as an audio file or text file.

[0912] A "database" is a storage system that allows the server to efficiently store, manage, and search conversation data received.

[0913] A "large-scale language model" is an algorithmic model for natural language processing that is trained on a large amount of text data.

[0914] "Training" refers to the process of feeding conversational data into a large-scale language model to learn specific speaking styles and phrases.

[0915] A "repository" is a storage system for storing trained conversational models.

[0916] "Profile information" refers to personal information and setting data that a user confirms after logging in to the system.

[0917] A "situation" is a specific situation or context in a conversation simulation, such as an interview format.

[0918] "Simulation" is the process of engaging in virtual interactions based on user-selected conversation partners and situations.

[0919] An "emotion engine" is an algorithm or system that recognizes emotions from user input data and provides the results to a server.

[0920] "Input data" refers to text and voice data provided by a user to a system.

[0921] "Emotional information" refers to data that indicates the emotional state of the user as recognized by the emotion engine.

[0922] "Log" refers to historical data that records interactions between users and conversation models.

[0923] "Optimization" refers to the process of appropriately adjusting the content and expression of a response based on recognized emotional information.

[0924] MODE FOR CARRYING OUT THE INVENTION

[0925] The present invention combines a system in which users learn languages ​​through realistic conversation simulations with an emotion engine to generate optimal responses according to the user's emotional state. The system includes the following main components:

[0926] 1. Terminal

[0927] A terminal is a device through which a data provider or user accesses the system. Examples include computers, smartphones, and tablets. Terminals are equipped with the following interfaces:

[0928] An interface for conversation data providers to upload conversation data

[0929] An interface for users to request conversational simulations and view the results of the simulations

[0930] 2. Server

[0931] The server receives, stores, cleans, and stores conversation data in a database. The server has the following functions:

[0932] 1. A function to receive conversation data and temporarily store it in storage

[0933] 2. A function to check the integrity of received data and store it in a database

[0934] 3. Ability to clean stored data and remove noise and private information

[0935] 4. The ability to convert cleaned data into a training set and feed it into a large-scale language model (LLM).

[0936] 5. Ability to save the conversation model generated through training and recall it whenever needed

[0937] 6. Ability to use conversational models to generate realistic responses based on user requests

[0938] 7. Ability to recognize user emotions using an emotion engine and generate optimal responses

[0939] 8. Ability to save conversation logs and use them as training data for future use

[0940] 3. Database

[0941] The database is a system that centrally stores and manages received conversation data and generated conversation models. Indexing enables efficient management and search.

[0942] 4. Emotion Engine

[0943] The emotion engine is an engine that recognizes emotions from user text input and voice data and optimizes responses based on those emotions.

[0944] Collection and processing of conversation data

[0945] Users use their devices to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. Audio files and text files can be used for this purpose. The device then sends the conversation data uploaded by the user to the server. The server checks the integrity of the received data and stores it in a database. Indexing allows for efficient management and searching.

[0946] Training the model

[0947] The server retrieves conversation data from the database and performs a data cleaning process. It converts the cleaned data into a training set and supplies it to the LLM. This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The server stores the generated conversation model in a repository and can be called up whenever needed.

[0948] Preparing for user interaction

[0949] A user logs in to the system using a terminal, checks their profile information, and selects a conversation partner and situation. The terminal then sends the selection to the server as a request. The server accepts the request, selects an appropriate conversation model from its saved storage, generates an initial response, and sends it to the terminal.

[0950] Real-time conversation simulation and emotion recognition

[0951] The user inputs the next question or dialogue phrase and sends it to the server via the device. The emotion engine analyzes the user's input data and recognizes emotions. The recognized emotion information is sent to the server. The server then optimizes the response based on the emotion information and sends the generated response to the device. This allows the user to effectively learn the language through realistic conversation simulation. The server saves these conversation logs and uses them as subsequent training data to improve the accuracy and quality of the system.

[0952] Specific examples

[0953] A user simulates an interview with a celebrity

[0954] 1. Terminal: The user logs in to the system and selects a celebrity to talk to. The situation is chosen to be an interview format.

[0955] 2. Server: Based on the user request, the celebrity conversation model is called and an initial response is generated, which reads, "Hello, celebrity. What would you like to talk about today?"

[0956] 3. Terminal: The initial response is displayed to the user. Type "Hello, celebrity. I'd like to ask you my next question. Tell me about your latest project."

[0957] 4. Emotion engine: Recognizes emotions from the user's input data and sends that information to the server. It is analyzed as "curious expression."

[0958] 5. Server: Based on the emotion information received from the emotion engine, the server generates and optimizes a response and sends it to the device. The response generated is "Thank you. Now, I'll tell you about my latest project..." and is displayed to the user.

[0959] 6. Server: Stores user interaction logs and uses them as training data for subsequent use to improve the accuracy of the system.

[0960] This allows users to advance their language learning while enjoying a realistic interview experience.

[0961] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0962] Specific processing steps of the system

[0963] Step 1:

[0964] A user logs into a system using an information terminal. The input is the user's authentication information (e.g., username, password), and the output is the start of a session for the authenticated user. To access the system, the user enters the appropriate authentication information. The system receives this, authenticates, and then starts the user session.

[0965] Step 2:

[0966] Users upload their own conversation data from their information terminals. The input is an audio file or a text file, and the output is the conversation data sent to the server. Using the system's interface, conversation data providers select a data file and upload it to the server.

[0967] Step 3:

[0968] The server temporarily stores the received conversation data in storage. The input is the conversation data sent from the terminal, and the output is the data temporarily stored in storage. After receiving the data, the server verifies it, performs an integrity check, and then stores the data in storage.

[0969] Step 4:

[0970] The server stores the saved conversation data in a database. The input is the data stored in temporary storage, and the output is the data indexed in the database. When storing the data in the database, an index is created to enable efficient management and search.

[0971] Step 5:

[0972] The server retrieves the conversation data from the database and runs a data cleaning process. The input is the raw data in the database, and the output is the cleaned data. The dataset is cleaned using algorithms to remove noise and privacy information.

[0973] Step 6:

[0974] The server converts the cleaned conversation data into a training set and feeds it to an LLM (large-scale language model). The input is the cleaned data, and the output is the training set. This completes the dataset for training the conversation model.

[0975] Step 7:

[0976] The server saves the conversation model generated by training in a repository. The input is the conversation model generated by training, and the output is the model saved in the repository. The generated model is saved in storage so that it can be used in response to any request.

[0977] Step 8:

[0978] The user selects a conversation partner and a situation using an information terminal. The input is the selection information of the conversation partner and the situation, and the output is request data to the server. The details of the conversation simulation are specified through the interface, and the information is sent to the system.

[0979] Step 9:

[0980] The server calls the appropriate stored conversation model based on the user's request and generates an initial response. The input is the user's request data, and the output is the generated initial response. The specified model is loaded and an initial response according to the request is generated.

[0981] Step 10:

[0982] The server sends the generated initial response to the information terminal. The input is the generated initial response, and the output is the response displayed on the terminal. The generated response is forwarded to the user for display.

[0983] Step 11:

[0984] The user inputs the next question or dialogue phrase and sends it to the server through the terminal. The input is the user's next question or dialogue phrase, and the output is the data sent to the server. The user inputs the next step to continue the dialogue.

[0985] Step 12:

[0986] The emotion engine analyzes the user's input data and recognizes emotions. The input is the user's next question or dialogue phrase, and the output is the recognized emotion information. It analyzes the provided data and recognizes the user's emotional state.

[0987] Step 13:

[0988] The emotion engine sends the recognized emotion information to the server. The input is the recognized emotion information, and the output is the data sent to the server. The analysis results are transferred to the server.

[0989] Step 14:

[0990] The server optimizes and generates a response based on the emotional information. The input is the emotional information and the user's next question, and the output is an optimized response. The response content is adjusted and generated according to the recognized emotion.

[0991] Step 15:

[0992] The server sends the generated response to the terminal and displays it to the user. The input is the generated response and the output is the response displayed on the terminal. The server forwards the response to the user and continues the dialogue.

[0993] Step 16:

[0994] The server stores the interaction between the user and the conversational model as a log. The input is the interaction data between the user and the conversational model, and the output is the stored log. The stored log is used for subsequent training.

[0995] (Application example 2)

[0996] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0997] Conventional language learning systems have had difficulty providing optimal responses based on conversational realism and emotions. It has also been difficult for users to engage in natural conversations in virtual stores while selecting products and receiving information in real time. This has led to problems such as reduced learning effectiveness and user satisfaction.

[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0999] In this invention, the server includes means for analyzing the emotional state of the user using an emotion engine and optimizing a response based on that information, means for providing the optimized response to the user's terminal in real time, and means for the user to select products and provide information in the virtual environment using the terminal, thereby enabling the user to obtain product information while engaging in natural conversation according to their emotions in real time in the virtual store.

[1000] A "conversation data provider" is a person or organization that uploads conversation data from a terminal.

[1001] "Terminal" means a device (e.g., computer, smartphone, tablet) used by a conversation data provider or user to access the system.

[1002] The "server" is the core computer system that receives and stores conversation data, serves it as training data, and ultimately generates responses.

[1003] A "database" is a system for storing and managing received conversation data and training data.

[1004] An "LLM (Large Scale Language Model)" is a language model trained using a large number of datasets to achieve natural language processing.

[1005] "Training" is the process of optimizing a large-scale language model using stored conversational data.

[1006] A "repository" is a system for storing and managing trained conversation models.

[1007] "User" means a person or entity that uses a terminal to access the system and use the conversation simulation.

[1008] An "emotion engine" is a technology that analyzes a user's emotional state and optimizes responses based on that information.

[1009] A "virtual environment" is a simulated environment that allows users to have a realistic experience within a digital space.

[1010] "Product selection" is the action of a user selecting a product of interest within a virtual environment.

[1011] "Information provision" means providing detailed information about the product selected by the user through a virtual assistant or the like.

[1012] "Real-time" means that processing is done immediately, without delay, and the results are provided to the user.

[1013] The following describes in detail an embodiment of the invention in a virtual store customer service system. First, the process of generating a program for the system that realizes this application example and its specific operation will be described.

[1014] Overall system configuration

[1015] The system consists of the following main elements:

[1016] 1. Terminal: A device through which a user accesses the system, such as a smartphone or head-mounted display.

[1017] 2. Server: The core computer system that receives, stores, trains, and generates responses from conversation data.

[1018] 3. Database: A system that stores and manages conversation data and conversation models.

[1019] 4. Emotion Engine: Technology that analyzes the user's emotional state and optimizes responses based on that information.

[1020] Detailed Operation

[1021] 1. User visits the virtual store:

[1022] By launching the application using a device (smartphone or head-mounted display), the user enters the 3D environment of the virtual store.

[1023] 2. Product Selection and Information:

[1024] Once the user selects the product of interest, product information is provided by the virtual assistant (trained conversation model).

[1025] 3. Real-time conversation simulation:

[1026] When a user asks a question to a virtual assistant, the emotion engine analyzes the question and the user's emotional state. The emotion engine's API (e.g., IBM Watson Tone Analyzer, Microsoft Azure Emotion API) analyzes the user's voice data and text to recognize the emotional state.

[1027] 4. Optimal response generation:

[1028] Based on the recognized emotional state, a large-scale language model (e.g., OpenAI's GPT series) generates an optimal response. The server sends a request to the triggered model and obtains the inference result.

[1029] 5. Response display:

[1030] The generated responses are displayed to the user through the device, and the conversation logs are saved by the server and used as training data for future use.

[1031] Specific hardware and software

[1032] Hardware

[1033] Smartphones: iPhone, Android, etc.

[1034] Head-mounted displays: Oculus Quest 2, HoloLens 2, etc.

[1035] software

[1036] Front-end frameworks: Unity, React Native, etc.

[1037] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.

[1038] Large-scale language models: OpenAI GPT-3, GPT-4, etc.

[1039] Database: MongoDB, MySQL, etc.

[1040] Server-side frameworks: Node.js, Django, etc.

[1041] Examples of concrete examples and prompts

[1042] Here's how the system works:

[1043] The user puts on a head-mounted display, enters a virtual store, and approaches specific products.

[1044] Example prompt: "What features does this new smartwatch have?"

[1045] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[1046] This allows users to obtain product information while engaging in natural, emotionally-responsive dialogue in real time in a virtual store.

[1047] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1048] Step 1:

[1049] Access to the system and entry to the virtual store

[1050] The user launches the application using a smartphone or head-mounted display, which accesses the 3D environment of the virtual store, where the user logs in to the system and the initial interface is displayed.

[1051] Input: User login information, device type

[1052] Data processing / calculation: User authentication, display of virtual store using 3D rendering engine

[1053] Output: Login success message, virtual store UI display

[1054] Step 2:

[1055] Product selection and initial information provision

[1056] The user moves through the virtual environment and selects a specific product. The device sends this selection information to the server, which retrieves the product's initial information from a database and displays it.

[1057] Input: ID of the product selected by the user

[1058] Data processing / calculation: database query for product information, information format conversion

[1059] Output: Product details (text, images, etc.)

[1060] Step 3:

[1061] Start the conversation simulation

[1062] The user inputs a question related to the product, and the device sends the question to the server, which then uses an emotion engine to analyze the user's emotional state.

[1063] Input: User text input (question content)

[1064] Data processing / calculation: Emotion analysis using emotion engines (e.g., IBM Watson Tone Analyzer)

[1065] Output: Emotional state data (e.g., surprise, joy, dissatisfaction)

[1066] Step 4:

[1067] Generating the best response

[1068] The server uses the emotional state data and the user's question to trigger a large-scale language model (e.g., OpenAI GPT-4) to generate an optimal response.

[1069] Input: User question, emotional state data

[1070] Data processing / calculation: Prompt generation and response generation for large-scale language models

[1071] Output: Optimized response text

[1072] Step 5:

[1073] Viewing the response

[1074] The generated response text is sent from the server to the terminal, which displays the response to the user.

[1075] Input: Response text data from the server

[1076] Data processing / calculation: rendering text data, updating the interface

[1077] Output: User confirms response text

[1078] Step 6:

[1079] Save conversation logs

[1080] The conversational exchanges are logged and stored by the server and used as training data for subsequent use.

[1081] Input: Conversation data between the user and the virtual assistant

[1082] Data processing / calculation: Logging to database, indexing

[1083] Output: Saved conversation log

[1084] Examples:

[1085] Prompt Sentence Examples

[1086] A user asks: "What features does this new smartwatch have?"

[1087] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[1088] This step allows users to receive appropriate responses in real time based on their emotions, enabling an efficient and satisfying virtual store experience.

[1089] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1090] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1091] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1092] [Third embodiment]

[1093] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1094] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1095] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1096] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1097] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1098] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1099] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1100] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1101] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1102] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1103] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1104] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1105] The present invention is a system that allows users to learn a language through realistic conversation simulations. This system provides simulations based on conversation data with famous people and professionals, and aims to improve learners' motivation and learning effectiveness.

[1106] System Configuration

[1107] The system includes the following major components:

[1108] 1. Terminal

[1109] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[1110] An interface for conversation data providers to upload conversation data.

[1111] An interface for users to request conversation simulations and view the simulation results.

[1112] 2. Server

[1113] The ability to receive, save, clean, and store conversation data in a database.

[1114] A function that trains LLMs (large-scale language models) based on saved conversation data.

[1115] Ability to save trained conversational models and recall them whenever needed.

[1116] The ability to use conversational models to generate realistic responses based on user requests.

[1117] A function that saves conversation logs and uses them as training data for future use.

[1118] 3. Database

[1119] A system that centrally stores and manages received conversation data and generated conversation models.

[1120] Collection and handling of conversation data

[1121] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[1122] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[1123] Training the model

[1124] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[1125] User Interaction

[1126] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[1127] Terminal: The selection is sent to the server as a request.

[1128] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[1129] Real-time conversation simulation

[1130] User: Enter the following conversation phrase and send it to the server via the device.

[1131] Server: Analyzes the user's input data, generates the next response, and sends it to the device. This interaction occurs in real time, allowing the user to simulate a realistic conversation experience with a real person.

[1132] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[1133] Specific examples

[1134] A user simulates an interview with celebrity B.

[1135] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[1136] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[1137] 3. Terminal: The initial response is displayed to the user.

[1138] 4. User: Enter the following question and submit it to the server:

[1139] 5. Server: Generates the following response and sends it to the device:

[1140] 6. Server: Stores the conversation log.

[1141] In this way, users can advance their language learning while simulating a real-life interview experience with celebrity B.

[1142] The processing flow will be explained below.

[1143] Step 1: Upload your conversation data

[1144] Terminal: Conversation data providers use terminals to upload conversation data (audio files or text files) to the system.

[1145] Server: Stores the received conversation data in temporary storage and checks the data integrity.

[1146] Step 2: Storing and indexing conversation data

[1147] Server: Stores the verified conversation data in a database.

[1148] Server: Creates an index of the stored data to facilitate future search and management.

[1149] Step 3: Cleaning the conversation data

[1150] Server: Cleans the stored conversation data, removing noise and privacy information.

[1151] Step 4: Generate a training set

[1152] Server: Converts the cleaned data into a training set for LLMs (large-scale language models).

[1153] Step 5: Train the model

[1154] Server: Feeds the training set to the LLM, allowing it to learn specific person speaking styles and phrases.

[1155] Server: Stores the trained model in a repository.

[1156] Step 6: User login and profile verification

[1157] User: Log in to the system and check your profile information.

[1158] Terminal: Sends the user's login information to the server for authentication.

[1159] Step 7: Choose your conversation partner and situation

[1160] User: Select a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the terminal.

[1161] Terminal: Sends the selection to the server as a request.

[1162] Step 8: Generate an initial response

[1163] Server: Based on the user's request, it invokes the appropriate conversation model to generate the initial response.

[1164] Server: Sends the generated initial response to the terminal.

[1165] Step 9: View the initial response

[1166] Terminal: Displays to the user the initial response received from the server.

[1167] Step 10: Enter conversation phrases

[1168] User: Enter the following conversation phrase into your device:

[1169] Terminal: Sends user input data to the server.

[1170] Step 11: Generate and Send Response

[1171] Server: Parses the user's input data and generates the following response:

[1172] Server: Sends the generated response to the terminal.

[1173] Step 12: Save the conversation log

[1174] Server: Conversation logs with the user are stored in a database and used as training data for future use.

[1175] Example 1

[1176] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1177] Language learning is becoming increasingly important in modern times, and learning through real conversations is particularly effective. However, opportunities to experience real conversations are limited, making it difficult to learn through conversation simulations with specific celebrities or professionals. Furthermore, systems that generate responses in real time in response to user requests are not common, making it difficult to provide a conversation experience that is close to reality. This has led to problems such as insufficient improvement in user motivation and effectiveness of learning.

[1178] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1179] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) and train it; a means for the server to store the trained conversation model in a repository; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response and send it to the terminal; a real-time conversation simulation means for the user to input the next conversation phrase and the server to generate a response and send it to the terminal; and a means for the server to store the interaction between the user and the conversation model as a log. This enables language learning through realistic conversation simulations with specific celebrities or professionals, thereby improving the motivation and effectiveness of users' learning.

[1180] "Conversation data provider" refers to a user or individual who provides conversation data to the conversation simulation system.

[1181] "Terminal" refers to a device, such as a computer, smartphone, or tablet, used by a conversation data provider or user to access the system.

[1182] "Conversation data" refers to materials such as audio files and text files used for conversation simulation.

[1183] "Server" refers to a computer system that receives, stores, processes data, trains models, etc.

[1184] "Database" refers to a system for efficiently storing, managing, and searching conversation data.

[1185] "LLM (Large-scale Language Model)" refers to an AI model for natural language processing that is trained based on massive amounts of conversational data.

[1186] "Training Set" refers to the collection of cleaned conversational data used to train the LLM.

[1187] "Repository" refers to a storage system for storing trained conversational models.

[1188] "User" refers to an individual who utilizes the conversation simulation system to simulate a real conversation experience.

[1189] "Situation" refers to the conversation format or setting selected by the user, such as an interview format.

[1190] A "request" refers to a request or instruction a user makes to a system.

[1191] "Response" means any reply or response content generated by LLM to a User's request.

[1192] "Real-time conversation simulation" refers to a process in which the server instantly generates and returns a response to a conversation phrase entered by the user.

[1193] "Conversation log" refers to data that records the interactions between a user and a conversation model.

[1194] "Cleaning" refers to the process of removing unnecessary noise and private information from conversation data.

[1195] The present invention is a system that allows users to learn a language through realistic conversation simulations. One feature of this system is that it provides simulations based on conversation data with famous people and professionals, thereby improving learners' motivation and learning effectiveness. An embodiment of the present invention is described in detail below.

[1196] This system mainly consists of three components: a terminal, a server, and a database.

[1197] Terminal

[1198] Terminals are devices through which conversation data providers and users access the system, and examples include computers, smartphones, and tablets. Conversation data providers use these terminals to access an interface for uploading conversation data such as audio files and text files to the system. Users can also request conversation simulations and view the simulation results through the terminals.

[1199] server

[1200] The server is responsible for the core processing of the system. First, it receives conversation data uploaded from the device and temporarily stores it in storage. Next, it checks the integrity of the data and stores it in a database. It also cleans the stored data to remove unnecessary noise and privacy information. The cleaned data is converted into a training set and fed into an LLM (large-scale language model). This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The generated conversation model is stored in a repository so that it can be called up when needed.

[1201] The server also invokes the trained conversation model based on the user's request to generate a realistic response. The generated response is sent to the device and displayed to the user. Furthermore, if the user enters the next conversational phrase, the server analyzes it, generates the next response, and sends it to the device. This interaction takes place in real time, allowing the user to simulate a realistic conversation experience with a real person. Finally, these conversation logs are saved and used as subsequent training data.

[1202] Database

[1203] The database is a system that centrally stores and manages received conversation data and generated conversation models. The database stores received conversation data with an index, allowing for efficient management and search.

[1204] Specific examples

[1205] Let us say that a user wants to simulate an interview with a well-known professional.

[1206] 1. User: Logs in to the system using a terminal, selects a "famous professional" as the conversation partner, and selects the "interview format" as the situation.

[1207] 2. Server: Based on the user request, it invokes the "famous professional" conversation model and generates an initial response, such as "Hello, what would you like to talk about today?"

[1208] 3. Terminal: The initial response generated is displayed to the user.

[1209] 4. User: Enters the following question: "What event has had the greatest impact on your career?" and submits it to the server.

[1210] 5. Server: Based on the received question, the following response is generated: "The event that had the greatest impact on my career is..." and sent to the device. This is done in real time.

[1211] An example prompt might look like this:

[1212] "You are interviewing a well-known professional. The first question is, 'What has been the most influential event in your career so far?' Please continue with the following interview content."

[1213] Through this simulation, users can learn languages ​​through realistic conversational experiences, improving learning motivation and effectiveness.

[1214] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1215] Step 1:

[1216] Uploading conversation data

[1217] Terminal: The conversation data provider uses the terminal to select conversation data such as interview content or dialogue scenes, and clicks the upload button. The input is an audio file (e.g., .mp3) or a text file (e.g., .txt). The output is an upload request.

[1218] Step 2:

[1219] Receiving and storing data

[1220] Server: Receives conversation data sent from the device and temporarily stores it in storage. The input is the upload request from the device, and the output is the temporarily stored conversation data. Specifically, this includes checking the file size and format to ensure there are no inconsistencies.

[1221] Step 3:

[1222] Data Integrity Check

[1223] Server: Checks the integrity of the received data. The input is the temporarily saved conversation data, and the output is the status of whether it can be stored in the database. Specific operations include checking the completeness and consistency of the data.

[1224] Step 4:

[1225] Data storage and indexing

[1226] Server: Stores the integrity-checked data in a database and creates an index. The input is the checked conversation data, and the output is an indexed database entry, allowing for efficient searches.

[1227] Step 5:

[1228] Cleaning the data

[1229] Server: Cleans the conversation data stored in the database and removes unnecessary noise and privacy information. The input is the conversation data retrieved from the database, and the output is the clean data. Specifically, noise filtering is performed using a text mining library.

[1230] Step 6:

[1231] Creating a training set

[1232] Server: Converts the cleaned conversation data into a training set and feeds it to the LLM (large-scale language model). The input is clean data, and the output is the training set. Specifically, it normalizes and tokenizes the data.

[1233] Step 7:

[1234] Training a conversation model

[1235] Server: Trains the LLM using the training set. The input is the training set, and the output is the trained conversation model. Specific operations include updating the model parameters.

[1236] Step 8:

[1237] Saving the conversation model

[1238] Server: Stores the trained conversational model in a repository. The input is the trained conversational model, and the output is the model stored in the repository, so that it can be recalled and used later.

[1239] Step 9:

[1240] Accepting user requests

[1241] User: Logs in to the system through a terminal, selects a conversation partner and a situation, and requests a simulation. The input is the user's selection information, and the output is the request transmission.

[1242] Step 10:

[1243] Generate an initial response

[1244] Server: Based on the user's request, it calls the appropriate conversation model and generates an initial response. The input is the user's request, and the output is the generated initial response. The specific process is to input a prompt to the conversation model and get a response.

[1245] Step 11:

[1246] Viewing the response

[1247] Terminal: Displays the generated initial response on the screen. The input is the initial response from the server, and the output is the response displayed to the user.

[1248] Step 12:

[1249] Enter the next conversation phrase

[1250] User: Receives the initial response, enters the next conversational phrase, and sends it to the server through the terminal. The input is the next conversational phrase, and the output is a request to the server.

[1251] Step 13:

[1252] Producing the following response

[1253] Server: Analyzes the user's input, generates the next response, and sends it to the device. The input is the user's next conversation phrase, and the output is the generated response. This process is done in real time.

[1254] Step 14:

[1255] Save conversation logs

[1256] Server: Stores the interactions between the user and the conversation model as a log. The input is the data of each conversation turn, and the output is the saved conversation log. The log is also used as training data for subsequent tasks.

[1257] (Application example 1)

[1258] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1259] Today's language learners have limited opportunities to learn languages ​​efficiently through real-life conversation experiences with specific experts or celebrities. This issue can decrease learners' motivation and reduce learning effectiveness. Furthermore, content distribution services, in particular, lack platforms where users can enjoy interactive content. This leads to issues such as reduced user engagement and a decrease in the appeal of the service.

[1260] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1261] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) for training; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the terminal; a means for the server to store the interaction between the user and the conversation model as a log; a means for the server to store the generated conversation log and provide it for later review or additional learning; and a means for the terminal to display responses to the user based on the selected conversation situation. This allows users to learn languages ​​through realistic conversation experiences with celebrities and experts, improving learning motivation and effectiveness. It also enables content distribution services to provide users with more interactive and engaging content.

[1262] A "conversation data provider" is an individual or organization that is responsible for uploading conversation data from a terminal.

[1263] "Terminal" refers to a device used by a conversation data provider or user to access the system, including a smartphone, tablet, head-mounted display, computer, etc.

[1264] "Server" is a computer system that receives, stores, and analyzes conversation data and manages and serves trained models.

[1265] The "database" is a system that centrally stores and manages received conversation data and generated conversation models.

[1266] An "LLM (Large-Scale Language Model)" is an artificial intelligence model trained on a huge amount of text data and capable of generating natural-sounding conversations.

[1267] A "repository" is a data storage system where trained conversational models are stored.

[1268] A "user" is an individual or group that accesses the system, selects a conversation partner and situation, and requests a simulation.

[1269] A "conversational model" is an AI model that is generated based on a trained LLM and is used to generate responses corresponding to specific conversational scenarios.

[1270] A "realistic conversational experience" is a process that simulates natural interactions that are close to real conversations.

[1271] A "conversation log" is a record of the interactions between a user and a conversation model, and is used as subsequent training data.

[1272] "Review" is the process of reviewing past conversation logs to enhance learning effectiveness.

[1273] The present invention provides a system for language learning that allows users to engage in realistic conversational simulations with celebrities and professionals, and is particularly suited to devices such as smartphones, tablets, and head-mounted displays.

[1274] System Configuration

[1275] 1. Hardware Configuration

[1276] Terminal: A device on which users can enjoy interactive conversation simulations. This includes smartphones, tablets, and head-mounted displays.

[1277] Server: A computer system that stores and manages conversation data, trains LLMs (large-scale language models), and generates responses.

[1278] Database: A system that centrally stores and manages conversation data and generated conversation models.

[1279] 2. Software Configuration

[1280] Flask: A microframework for building web applications that handles server-side processing.

[1281] OpenAI API: Used to generate natural-sounding conversational responses based on user input using generative AI models.

[1282] Data processing and calculation

[1283] Receiving user input: The server receives phrases entered by the user on the device, including text input and voice input.

[1284] Prompt Generation: Based on the user's input, generate a prompt to continue the virtual conversation, for example, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1285] Response Generation: Uses OpenAI's ChatGPT API to generate appropriate responses based on the generated prompts.

[1286] Returning the response: The generated response is sent to the terminal and displayed to the user.

[1287] Saving conversation logs: Records of interactions between the user and the conversation model (conversation logs) are saved on the server.

[1288] Specific examples

[1289] If a user is trying to learn Japanese, the process would be as follows:

[1290] 1. User: I want to simulate a conversation with a famous actor.

[1291] 2. User Input: "Tell me about your passion for filmmaking."

[1292] 3. Server: Takes input and generates the following prompt:

[1293] Prompt: "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1294] 4. OpenAI API: Generate appropriate responses based on prompts.

[1295] 5. Terminal: Displays the generated response to the user.

[1296] In this way, users can effectively learn languages ​​through realistic conversation experiences with celebrities and experts.

[1297] The above is the "Mode for Carrying Out the Invention."

[1298] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1299] Step 1:

[1300] The user selects a conversation partner and situation from the terminal and requests a simulation.

[1301] Input: User selection of conversation partner and situation.

[1302] Output: The request data to the server.

[1303] How it works: A user logs into the application using a smartphone, tablet, or head-mounted display, selects a conversation partner (e.g., a famous actor) and a situation for language learning, and then sends the selection to the server.

[1304] Step 2:

[1305] The server receives the request and invokes the conversation model.

[1306] Input: Request data from the user.

[1307] Output: The initial response of the conversation model.

[1308] Specific operation: Based on the received request, the server retrieves the corresponding conversation model from the database, for example, selects the conversation model of a famous actor, and generates an initial response.

[1309] Step 3:

[1310] The server receives the user's input and generates a prompt.

[1311] Input: A phrase entered by the user through the device (e.g., "Tell me about your passion for filmmaking.").

[1312] Output: The generated prompt statement.

[1313] Specific behavior: Based on the user's input phrase, the server generates a prompt sentence. The generated prompt sentence will be something like, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1314] Step 4:

[1315] The server calls the OpenAI API to generate a response based on the prompt.

[1316] Input: The generated prompt statement.

[1317] Output: The generated response.

[1318] Specific operation: The server uses the OpenAI API to generate an appropriate response based on the generated prompt. At this stage, the prompt is sent to the API and the response text is retrieved.

[1319] Step 5:

[1320] The server generates a response that is sent to the terminal and displayed to the user.

[1321] Input: The response obtained from the OpenAI API.

[1322] Output: The response text that is displayed on the terminal.

[1323] Specific operation: The server sends the generated response to the terminal, and the user confirms this response through the terminal. For example, the response displayed is "About your passion for filmmaking."

[1324] Step 6:

[1325] The server stores the interaction between the user and the conversation model as a log.

[1326] Input: The dialogue between the user and the conversation model.

[1327] Output: Saved conversation logs.

[1328] How it works: The server stores the interactions between the user and the conversation model in a database for subsequent learning and review. Conversation logs are automatically saved and made available for the user to refer to later.

[1329] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1330] The present invention combines an emotion engine with a language learning system that allows users to learn languages ​​through realistic conversation simulations, generating optimal responses according to the user's emotional state. This system aims to improve learners' motivation and learning effectiveness by providing simulations based on conversation data with celebrities and professionals, in particular.

[1331] System Configuration

[1332] The system includes the following major components:

[1333] 1. Terminal

[1334] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[1335] An interface for conversation data providers to upload conversation data.

[1336] An interface for users to request conversation simulations and view the simulation results.

[1337] 2. Server

[1338] The ability to receive, save, clean, and store conversation data in a database.

[1339] A function that trains LLMs (large-scale language models) based on saved conversation data.

[1340] Ability to save trained conversational models and recall them whenever needed.

[1341] The ability to use conversational models to generate realistic responses based on user requests.

[1342] Ability to recognize user emotions using an emotion engine and generate optimal responses.

[1343] A function that saves conversation logs and uses them as training data for future use.

[1344] 3. Database

[1345] A system that centrally stores and manages received conversation data and generated conversation models.

[1346] 4. Emotion Engine

[1347] An engine for recognizing emotions from user text input and voice data and optimizing responses based on that.

[1348] Collection and handling of conversation data

[1349] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[1350] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[1351] Training the model

[1352] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[1353] User Interaction

[1354] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[1355] Terminal: The selection is sent to the server as a request.

[1356] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[1357] Real-time conversation simulation and emotion recognition

[1358] Emotion engine: Recognizes emotions from the user's text input or voice data. The recognized emotion information is sent to the server.

[1359] Server: Optimizes responses based on the emotional information received from the emotion engine. Sends the generated responses to the device so that the conversation continues in real time.

[1360] User: Enter the following conversation phrase and send it to the server via the device. If an emotion is recognized, that information is also sent.

[1361] Server: Analyzes the user's input data and emotional information, and generates the following response, which is also sent to the device in an optimized form.

[1362] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[1363] Specific examples

[1364] A user simulates an interview with celebrity B.

[1365] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[1366] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[1367] 3. Terminal: The initial response is displayed to the user.

[1368] 4. User: Enter the following question and send it to the server via the terminal.

[1369] 5. Emotion engine: Recognizes emotions from the user's input data and also sends this information to the server.

[1370] 6. Server: Based on the emotion information received from the emotion engine, it generates and optimizes the next response and sends it to the terminal.

[1371] 7. Server: Saves the conversation log.

[1372] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[1373] The processing flow will be explained below.

[1374] Step 1: Upload your conversation data

[1375] Terminal: Conversation data providers use terminals to upload conversation data (voice or text) to the system.

[1376] Server: Receives conversation data and temporarily stores it in storage.

[1377] Step 2: Verify the integrity of the conversation data and store it

[1378] Server: Checks the integrity of the received conversation data and stores it in a database, adding appropriate metadata when storing it.

[1379] Step 3: Indexing

[1380] Server: Indexes the stored conversation data, enabling efficient search and management.

[1381] Step 4: Data Cleaning

[1382] Server: Cleans the stored conversation data to remove noise and privacy information.

[1383] Step 5: Generate a training set

[1384] Server: Converts the cleaned data into a training set and feeds it into the LLM (large-scale language model).

[1385] Step 6: Train and save the model

[1386] Server: Trains the LLM using the training set to learn the speaking style and phrases of a particular person.

[1387] Server: Stores the trained conversation model in a repository.

[1388] Step 7: Logging in users

[1389] User: Log in to the system and check your profile information.

[1390] Terminal: Sends login information to the server for authentication.

[1391] Step 8: Choose your conversation partner and situation

[1392] User: Selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the system.

[1393] Terminal: Sends the selection to the server.

[1394] Step 9: Receiving a request and generating an initial response

[1395] Server: Receives requests from users, invokes the selected conversation model, and generates the initial response.

[1396] Server: Sends the generated initial response to the terminal.

[1397] Step 10: View the initial response

[1398] Terminal: Displays the initial response received from the server to the user.

[1399] Step 11: Enter conversation phrases

[1400] User: Enter the following conversation phrase into your device:

[1401] Terminal: Sends the entered data to the server.

[1402] Step 12: Recognize emotions

[1403] Server: Uses an emotion engine to recognize emotions from user input data and voice data.

[1404] Step 13: Generate and optimize responses

[1405] Server: Generates the optimal response based on the user's input data and recognized emotional information.

[1406] Server: Sends the generated optimization response to the terminal.

[1407] Step 14: View the response

[1408] Terminal: displays the optimization response received from the server to the user.

[1409] Step 15: Save the conversation log

[1410] Server: Saves conversation logs with users in a database and uses them as training data for future use.

[1411] Specific examples

[1412] A user simulates an interview with celebrity B.

[1413] 1. Step 1:

[1414] Device: Celebrity B uploads his / her interview recording.

[1415] Server: Receives and temporarily stores conversation data.

[1416] 2. Step 2:

[1417] Server: Checks the integrity of the data and stores it in the database.

[1418] 3. Step 3:

[1419] Server: Creates indexes to simplify later searches.

[1420] 4. Steps 4-6:

[1421] Server: Performs data cleaning, generates training sets, and trains the LLM.

[1422] 5. Steps 7-8:

[1423] User: Log in to the system and select Celebrity B. Select the interview format.

[1424] Terminal: Sends the request contents to the server.

[1425] 6. Steps 9-10:

[1426] Server: Generates an initial response and sends it to the terminal.

[1427] Terminal: Display the initial response to the user.

[1428] 7. Steps 11-12:

[1429] User: Enter the following question and send it to the server via the terminal.

[1430] Server: Recognizes emotions using an emotion engine.

[1431] 8. Steps 13-14:

[1432] Server: Generates the optimal response based on the recognized emotional information and sends it to the device.

[1433] Terminal: Display the optimization response to the user.

[1434] 9. Step 15:

[1435] Server: Saves conversation logs and uses them as training data for future use.

[1436] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[1437] Example 2

[1438] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1439] Existing language learning systems have difficulty enabling users to effectively learn a language through realistic conversation simulations. Furthermore, they lack a mechanism for generating optimal responses based on the user's emotional state. As a result, learners' motivation and learning effectiveness may decrease.

[1440] The specification processing by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a conversation data provider to upload conversation data from an information terminal; a means for the server to save the received conversation data in a database; a means for the server to supply the saved conversation data to a large-scale language model and train it; a means for the server to save the trained conversation model in a repository; a means for a user to select a conversation partner and a situation from an information terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the information terminal; a means for an emotion engine to recognize emotions from the user's input data and send the emotion information to the server; a means for the server to optimize and generate a response based on the emotion information and send it to the information terminal; and a means for the server to save the interaction between the user and the conversation model as a log. This enables users to effectively learn language through realistic conversation simulations and improve their learning experience.

[1441] "Conversation Data Provider" means a person or organization that provides conversation data to the system.

[1442] An "information terminal" is an electronic device used to access a system and send and receive data. Examples include computers, smartphones, and tablets.

[1443] A "server" is a computer system responsible for receiving, processing, and storing conversation data.

[1444] "Conversation Data" refers to a record of a conversation provided as an audio file or text file.

[1445] A "database" is a storage system that allows the server to efficiently store, manage, and search conversation data received.

[1446] A "large-scale language model" is an algorithmic model for natural language processing that is trained on a large amount of text data.

[1447] "Training" refers to the process of feeding conversational data into a large-scale language model to learn specific speaking styles and phrases.

[1448] A "repository" is a storage system for storing trained conversational models.

[1449] "Profile information" refers to personal information and setting data that a user confirms after logging in to the system.

[1450] A "situation" is a specific situation or context in a conversation simulation, such as an interview format.

[1451] "Simulation" is the process of engaging in virtual interactions based on user-selected conversation partners and situations.

[1452] An "emotion engine" is an algorithm or system that recognizes emotions from user input data and provides the results to a server.

[1453] "Input data" refers to text and voice data provided by a user to a system.

[1454] "Emotional information" refers to data that indicates the emotional state of the user as recognized by the emotion engine.

[1455] "Log" refers to historical data that records interactions between users and conversation models.

[1456] "Optimization" refers to the process of appropriately adjusting the content and expression of a response based on recognized emotional information.

[1457] MODE FOR CARRYING OUT THE INVENTION

[1458] The present invention combines a system in which users learn languages ​​through realistic conversation simulations with an emotion engine to generate optimal responses according to the user's emotional state. The system includes the following main components:

[1459] 1. Terminal

[1460] A terminal is a device through which a data provider or user accesses the system. Examples include computers, smartphones, and tablets. Terminals are equipped with the following interfaces:

[1461] An interface for conversation data providers to upload conversation data

[1462] An interface for users to request conversational simulations and view the results of the simulations

[1463] 2. Server

[1464] The server receives, stores, cleans, and stores conversation data in a database. The server has the following functions:

[1465] 1. A function to receive conversation data and temporarily store it in storage

[1466] 2. A function to check the integrity of received data and store it in a database

[1467] 3. Ability to clean stored data and remove noise and private information

[1468] 4. The ability to convert cleaned data into a training set and feed it into a large-scale language model (LLM).

[1469] 5. Ability to save the conversation model generated through training and recall it whenever needed

[1470] 6. Ability to use conversational models to generate realistic responses based on user requests

[1471] 7. Ability to recognize user emotions using an emotion engine and generate optimal responses

[1472] 8. Ability to save conversation logs and use them as training data for future use

[1473] 3. Database

[1474] The database is a system that centrally stores and manages received conversation data and generated conversation models. Indexing enables efficient management and search.

[1475] 4. Emotion Engine

[1476] The emotion engine is an engine that recognizes emotions from user text input and voice data and optimizes responses based on those emotions.

[1477] Collection and processing of conversation data

[1478] Users use their devices to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. Audio files and text files can be used for this purpose. The device then sends the conversation data uploaded by the user to the server. The server checks the integrity of the received data and stores it in a database. Indexing allows for efficient management and searching.

[1479] Training the model

[1480] The server retrieves conversation data from the database and performs a data cleaning process. It converts the cleaned data into a training set and supplies it to the LLM. This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The server stores the generated conversation model in a repository and can be called up whenever needed.

[1481] Preparing for user interaction

[1482] A user logs in to the system using a terminal, checks their profile information, and selects a conversation partner and situation. The terminal then sends the selection to the server as a request. The server accepts the request, selects an appropriate conversation model from its saved storage, generates an initial response, and sends it to the terminal.

[1483] Real-time conversation simulation and emotion recognition

[1484] The user inputs the next question or dialogue phrase and sends it to the server via the device. The emotion engine analyzes the user's input data and recognizes emotions. The recognized emotion information is sent to the server. The server then optimizes the response based on the emotion information and sends the generated response to the device. This allows the user to effectively learn the language through realistic conversation simulation. The server saves these conversation logs and uses them as subsequent training data to improve the accuracy and quality of the system.

[1485] Specific examples

[1486] A user simulates an interview with a celebrity

[1487] 1. Terminal: The user logs in to the system and selects a celebrity to talk to. The situation is chosen to be an interview format.

[1488] 2. Server: Based on the user request, the celebrity conversation model is called and an initial response is generated, which reads, "Hello, celebrity. What would you like to talk about today?"

[1489] 3. Terminal: The initial response is displayed to the user. Type "Hello, celebrity. I'd like to ask you my next question. Tell me about your latest project."

[1490] 4. Emotion engine: Recognizes emotions from the user's input data and sends that information to the server. It is analyzed as "curious expression."

[1491] 5. Server: Based on the emotion information received from the emotion engine, the server generates and optimizes a response and sends it to the device. The response generated is "Thank you. Now, I'll tell you about my latest project..." and is displayed to the user.

[1492] 6. Server: Stores user interaction logs and uses them as training data for subsequent use to improve the accuracy of the system.

[1493] This allows users to advance their language learning while enjoying a realistic interview experience.

[1494] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1495] Specific processing steps of the system

[1496] Step 1:

[1497] A user logs into a system using an information terminal. The input is the user's authentication information (e.g., username, password), and the output is the start of a session for the authenticated user. To access the system, the user enters the appropriate authentication information. The system receives this, authenticates, and then starts the user session.

[1498] Step 2:

[1499] Users upload their own conversation data from their information terminals. The input is an audio file or a text file, and the output is the conversation data sent to the server. Using the system's interface, conversation data providers select a data file and upload it to the server.

[1500] Step 3:

[1501] The server temporarily stores the received conversation data in storage. The input is the conversation data sent from the terminal, and the output is the data temporarily stored in storage. After receiving the data, the server verifies it, performs an integrity check, and then stores the data in storage.

[1502] Step 4:

[1503] The server stores the saved conversation data in a database. The input is the data stored in temporary storage, and the output is the data indexed in the database. When storing the data in the database, an index is created to enable efficient management and search.

[1504] Step 5:

[1505] The server retrieves the conversation data from the database and runs a data cleaning process. The input is the raw data in the database, and the output is the cleaned data. The dataset is cleaned using algorithms to remove noise and privacy information.

[1506] Step 6:

[1507] The server converts the cleaned conversation data into a training set and feeds it to an LLM (large-scale language model). The input is the cleaned data, and the output is the training set. This completes the dataset for training the conversation model.

[1508] Step 7:

[1509] The server saves the conversation model generated by training in a repository. The input is the conversation model generated by training, and the output is the model saved in the repository. The generated model is saved in storage so that it can be used in response to any request.

[1510] Step 8:

[1511] The user selects a conversation partner and a situation using an information terminal. The input is the selection information of the conversation partner and the situation, and the output is request data to the server. The details of the conversation simulation are specified through the interface, and the information is sent to the system.

[1512] Step 9:

[1513] The server calls the appropriate stored conversation model based on the user's request and generates an initial response. The input is the user's request data, and the output is the generated initial response. The specified model is loaded and an initial response according to the request is generated.

[1514] Step 10:

[1515] The server sends the generated initial response to the information terminal. The input is the generated initial response, and the output is the response displayed on the terminal. The generated response is forwarded to the user for display.

[1516] Step 11:

[1517] The user inputs the next question or dialogue phrase and sends it to the server through the terminal. The input is the user's next question or dialogue phrase, and the output is the data sent to the server. The user inputs the next step to continue the dialogue.

[1518] Step 12:

[1519] The emotion engine analyzes the user's input data and recognizes emotions. The input is the user's next question or dialogue phrase, and the output is the recognized emotion information. It analyzes the provided data and recognizes the user's emotional state.

[1520] Step 13:

[1521] The emotion engine sends the recognized emotion information to the server. The input is the recognized emotion information, and the output is the data sent to the server. The analysis results are transferred to the server.

[1522] Step 14:

[1523] The server optimizes and generates a response based on the emotional information. The input is the emotional information and the user's next question, and the output is an optimized response. The response content is adjusted and generated according to the recognized emotion.

[1524] Step 15:

[1525] The server sends the generated response to the terminal and displays it to the user. The input is the generated response and the output is the response displayed on the terminal. The server forwards the response to the user and continues the dialogue.

[1526] Step 16:

[1527] The server stores the interaction between the user and the conversational model as a log. The input is the interaction data between the user and the conversational model, and the output is the stored log. The stored log is used for subsequent training.

[1528] (Application example 2)

[1529] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1530] Conventional language learning systems have had difficulty providing optimal responses based on conversational realism and emotions. It has also been difficult for users to engage in natural conversations in virtual stores while selecting products and receiving information in real time. This has led to problems such as reduced learning effectiveness and user satisfaction.

[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1532] In this invention, the server includes means for analyzing the emotional state of the user using an emotion engine and optimizing a response based on that information, means for providing the optimized response to the user's terminal in real time, and means for the user to select products and provide information in the virtual environment using the terminal, thereby enabling the user to obtain product information while engaging in natural conversation according to their emotions in real time in the virtual store.

[1533] A "conversation data provider" is a person or organization that uploads conversation data from a terminal.

[1534] "Terminal" means a device (e.g., computer, smartphone, tablet) used by a conversation data provider or user to access the system.

[1535] The "server" is the core computer system that receives and stores conversation data, serves it as training data, and ultimately generates responses.

[1536] A "database" is a system for storing and managing received conversation data and training data.

[1537] An "LLM (Large Scale Language Model)" is a language model trained using a large number of datasets to achieve natural language processing.

[1538] "Training" is the process of optimizing a large-scale language model using stored conversational data.

[1539] A "repository" is a system for storing and managing trained conversation models.

[1540] "User" means a person or entity that uses a terminal to access the system and use the conversation simulation.

[1541] An "emotion engine" is a technology that analyzes a user's emotional state and optimizes responses based on that information.

[1542] A "virtual environment" is a simulated environment that allows users to have a realistic experience within a digital space.

[1543] "Product selection" is the action of a user selecting a product of interest within a virtual environment.

[1544] "Information provision" means providing detailed information about the product selected by the user through a virtual assistant or the like.

[1545] "Real-time" means that processing is done immediately, without delay, and the results are provided to the user.

[1546] The following describes in detail an embodiment of the invention in a virtual store customer service system. First, the process of generating a program for the system that realizes this application example and its specific operation will be described.

[1547] Overall system configuration

[1548] The system consists of the following main elements:

[1549] 1. Terminal: A device through which a user accesses the system, such as a smartphone or head-mounted display.

[1550] 2. Server: The core computer system that receives, stores, trains, and generates responses from conversation data.

[1551] 3. Database: A system that stores and manages conversation data and conversation models.

[1552] 4. Emotion Engine: Technology that analyzes the user's emotional state and optimizes responses based on that information.

[1553] Detailed Operation

[1554] 1. User visits the virtual store:

[1555] By launching the application using a device (smartphone or head-mounted display), the user enters the 3D environment of the virtual store.

[1556] 2. Product Selection and Information:

[1557] Once the user selects the product of interest, product information is provided by the virtual assistant (trained conversation model).

[1558] 3. Real-time conversation simulation:

[1559] When a user asks a question to a virtual assistant, the emotion engine analyzes the question and the user's emotional state. The emotion engine's API (e.g., IBM Watson Tone Analyzer, Microsoft Azure Emotion API) analyzes the user's voice data and text to recognize the emotional state.

[1560] 4. Optimal response generation:

[1561] Based on the recognized emotional state, a large-scale language model (e.g., OpenAI's GPT series) generates an optimal response. The server sends a request to the triggered model and obtains the inference result.

[1562] 5. Response display:

[1563] The generated responses are displayed to the user through the device, and the conversation logs are saved by the server and used as training data for future use.

[1564] Specific hardware and software

[1565] Hardware

[1566] Smartphones: iPhone, Android, etc.

[1567] Head-mounted displays: Oculus Quest 2, HoloLens 2, etc.

[1568] software

[1569] Front-end frameworks: Unity, React Native, etc.

[1570] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.

[1571] Large-scale language models: OpenAI GPT-3, GPT-4, etc.

[1572] Database: MongoDB, MySQL, etc.

[1573] Server-side frameworks: Node.js, Django, etc.

[1574] Examples of concrete examples and prompts

[1575] Here's how the system works:

[1576] The user puts on a head-mounted display, enters a virtual store, and approaches specific products.

[1577] Example prompt: "What features does this new smartwatch have?"

[1578] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[1579] This allows users to obtain product information while engaging in natural, emotionally-responsive dialogue in real time in a virtual store.

[1580] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1581] Step 1:

[1582] Access to the system and entry to the virtual store

[1583] The user launches the application using a smartphone or head-mounted display, which accesses the 3D environment of the virtual store, where the user logs in to the system and the initial interface is displayed.

[1584] Input: User login information, device type

[1585] Data processing / calculation: User authentication, display of virtual store using 3D rendering engine

[1586] Output: Login success message, virtual store UI display

[1587] Step 2:

[1588] Product selection and initial information provision

[1589] The user moves through the virtual environment and selects a specific product. The device sends this selection information to the server, which retrieves the product's initial information from a database and displays it.

[1590] Input: ID of the product selected by the user

[1591] Data processing / calculation: database query for product information, information format conversion

[1592] Output: Product details (text, images, etc.)

[1593] Step 3:

[1594] Start the conversation simulation

[1595] The user inputs a question related to the product, and the device sends the question to the server, which then uses an emotion engine to analyze the user's emotional state.

[1596] Input: User text input (question content)

[1597] Data processing / calculation: Emotion analysis using emotion engines (e.g., IBM Watson Tone Analyzer)

[1598] Output: Emotional state data (e.g., surprise, joy, dissatisfaction)

[1599] Step 4:

[1600] Generating the best response

[1601] The server uses the emotional state data and the user's question to trigger a large-scale language model (e.g., OpenAI GPT-4) to generate an optimal response.

[1602] Input: User question, emotional state data

[1603] Data processing / calculation: Prompt generation and response generation for large-scale language models

[1604] Output: Optimized response text

[1605] Step 5:

[1606] Viewing the response

[1607] The generated response text is sent from the server to the terminal, which displays the response to the user.

[1608] Input: Response text data from the server

[1609] Data processing / calculation: rendering text data, updating the interface

[1610] Output: User confirms response text

[1611] Step 6:

[1612] Save conversation logs

[1613] The conversational exchanges are logged and stored by the server and used as training data for subsequent use.

[1614] Input: Conversation data between the user and the virtual assistant

[1615] Data processing / calculation: Logging to database, indexing

[1616] Output: Saved conversation log

[1617] Examples:

[1618] Prompt Sentence Examples

[1619] A user asks: "What features does this new smartwatch have?"

[1620] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[1621] This step allows users to receive appropriate responses in real time based on their emotions, enabling an efficient and satisfying virtual store experience.

[1622] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1623] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1624] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1625] [Fourth embodiment]

[1626] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1627] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1628] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1629] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1630] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1631] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1632] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1633] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1634] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1635] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1636] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1637] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1638] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1639] The present invention is a system that allows users to learn a language through realistic conversation simulations. This system provides simulations based on conversation data with famous people and professionals, and aims to improve learners' motivation and learning effectiveness.

[1640] System Configuration

[1641] The system includes the following major components:

[1642] 1. Terminal

[1643] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[1644] An interface for conversation data providers to upload conversation data.

[1645] An interface for users to request conversation simulations and view the simulation results.

[1646] 2. Server

[1647] The ability to receive, save, clean, and store conversation data in a database.

[1648] A function that trains LLMs (large-scale language models) based on saved conversation data.

[1649] Ability to save trained conversational models and recall them whenever needed.

[1650] The ability to use conversational models to generate realistic responses based on user requests.

[1651] A function that saves conversation logs and uses them as training data for future use.

[1652] 3. Database

[1653] A system that centrally stores and manages received conversation data and generated conversation models.

[1654] Collection and handling of conversation data

[1655] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[1656] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[1657] Training the model

[1658] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[1659] User Interaction

[1660] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[1661] Terminal: The selection is sent to the server as a request.

[1662] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[1663] Real-time conversation simulation

[1664] User: Enter the following conversation phrase and send it to the server via the device.

[1665] Server: Analyzes the user's input data, generates the next response, and sends it to the device. This interaction occurs in real time, allowing the user to simulate a realistic conversation experience with a real person.

[1666] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[1667] Specific examples

[1668] A user simulates an interview with celebrity B.

[1669] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[1670] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[1671] 3. Terminal: The initial response is displayed to the user.

[1672] 4. User: Enter the following question and submit it to the server:

[1673] 5. Server: Generates the following response and sends it to the device:

[1674] 6. Server: Stores the conversation log.

[1675] In this way, users can advance their language learning while simulating a real-life interview experience with celebrity B.

[1676] The processing flow will be explained below.

[1677] Step 1: Upload your conversation data

[1678] Terminal: Conversation data providers use terminals to upload conversation data (audio files or text files) to the system.

[1679] Server: Stores the received conversation data in temporary storage and checks the data integrity.

[1680] Step 2: Storing and indexing conversation data

[1681] Server: Stores the verified conversation data in a database.

[1682] Server: Creates an index of the stored data to facilitate future search and management.

[1683] Step 3: Cleaning the conversation data

[1684] Server: Cleans the stored conversation data, removing noise and privacy information.

[1685] Step 4: Generate a training set

[1686] Server: Converts the cleaned data into a training set for LLMs (large-scale language models).

[1687] Step 5: Train the model

[1688] Server: Feeds the training set to the LLM, allowing it to learn specific person speaking styles and phrases.

[1689] Server: Stores the trained model in a repository.

[1690] Step 6: User login and profile verification

[1691] User: Log in to the system and check your profile information.

[1692] Terminal: Sends the user's login information to the server for authentication.

[1693] Step 7: Choose your conversation partner and situation

[1694] User: Select a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the terminal.

[1695] Terminal: Sends the selection to the server as a request.

[1696] Step 8: Generate an initial response

[1697] Server: Based on the user's request, it invokes the appropriate conversation model to generate the initial response.

[1698] Server: Sends the generated initial response to the terminal.

[1699] Step 9: View the initial response

[1700] Terminal: Displays to the user the initial response received from the server.

[1701] Step 10: Enter conversation phrases

[1702] User: Enter the following conversation phrase into your device:

[1703] Terminal: Sends user input data to the server.

[1704] Step 11: Generate and Send Response

[1705] Server: Parses the user's input data and generates the following response:

[1706] Server: Sends the generated response to the terminal.

[1707] Step 12: Save the conversation log

[1708] Server: Conversation logs with the user are stored in a database and used as training data for future use.

[1709] Example 1

[1710] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1711] Language learning is becoming increasingly important in modern times, and learning through real conversations is particularly effective. However, opportunities to experience real conversations are limited, making it difficult to learn through conversation simulations with specific celebrities or professionals. Furthermore, systems that generate responses in real time in response to user requests are not common, making it difficult to provide a conversation experience that is close to reality. This has led to problems such as insufficient improvement in user motivation and effectiveness of learning.

[1712] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1713] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) and train it; a means for the server to store the trained conversation model in a repository; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response and send it to the terminal; a real-time conversation simulation means for the user to input the next conversation phrase and the server to generate a response and send it to the terminal; and a means for the server to store the interaction between the user and the conversation model as a log. This enables language learning through realistic conversation simulations with specific celebrities or professionals, thereby improving the motivation and effectiveness of users' learning.

[1714] "Conversation data provider" refers to a user or individual who provides conversation data to the conversation simulation system.

[1715] "Terminal" refers to a device, such as a computer, smartphone, or tablet, used by a conversation data provider or user to access the system.

[1716] "Conversation data" refers to materials such as audio files and text files used for conversation simulation.

[1717] "Server" refers to a computer system that receives, stores, processes data, trains models, etc.

[1718] "Database" refers to a system for efficiently storing, managing, and searching conversation data.

[1719] "LLM (Large-scale Language Model)" refers to an AI model for natural language processing that is trained based on massive amounts of conversational data.

[1720] "Training Set" refers to the collection of cleaned conversational data used to train the LLM.

[1721] "Repository" refers to a storage system for storing trained conversational models.

[1722] "User" refers to an individual who utilizes the conversation simulation system to simulate a real conversation experience.

[1723] "Situation" refers to the conversation format or setting selected by the user, such as an interview format.

[1724] A "request" refers to a request or instruction a user makes to a system.

[1725] "Response" means any reply or response content generated by LLM to a User's request.

[1726] "Real-time conversation simulation" refers to a process in which the server instantly generates and returns a response to a conversation phrase entered by the user.

[1727] "Conversation log" refers to data that records the interactions between a user and a conversation model.

[1728] "Cleaning" refers to the process of removing unnecessary noise and private information from conversation data.

[1729] The present invention is a system that allows users to learn a language through realistic conversation simulations. One feature of this system is that it provides simulations based on conversation data with famous people and professionals, thereby improving learners' motivation and learning effectiveness. An embodiment of the present invention is described in detail below.

[1730] This system mainly consists of three components: a terminal, a server, and a database.

[1731] Terminal

[1732] Terminals are devices through which conversation data providers and users access the system, and examples include computers, smartphones, and tablets. Conversation data providers use these terminals to access an interface for uploading conversation data such as audio files and text files to the system. Users can also request conversation simulations and view the simulation results through the terminals.

[1733] server

[1734] The server is responsible for the core processing of the system. First, it receives conversation data uploaded from the device and temporarily stores it in storage. Next, it checks the integrity of the data and stores it in a database. It also cleans the stored data to remove unnecessary noise and privacy information. The cleaned data is converted into a training set and fed into an LLM (large-scale language model). This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The generated conversation model is stored in a repository so that it can be called up when needed.

[1735] The server also invokes the trained conversation model based on the user's request to generate a realistic response. The generated response is sent to the device and displayed to the user. Furthermore, if the user enters the next conversational phrase, the server analyzes it, generates the next response, and sends it to the device. This interaction takes place in real time, allowing the user to simulate a realistic conversation experience with a real person. Finally, these conversation logs are saved and used as subsequent training data.

[1736] Database

[1737] The database is a system that centrally stores and manages received conversation data and generated conversation models. The database stores received conversation data with an index, allowing for efficient management and search.

[1738] Specific examples

[1739] Let us say that a user wants to simulate an interview with a well-known professional.

[1740] 1. User: Logs in to the system using a terminal, selects a "famous professional" as the conversation partner, and selects the "interview format" as the situation.

[1741] 2. Server: Based on the user request, it invokes the "famous professional" conversation model and generates an initial response, such as "Hello, what would you like to talk about today?"

[1742] 3. Terminal: The initial response generated is displayed to the user.

[1743] 4. User: Enters the following question: "What event has had the greatest impact on your career?" and submits it to the server.

[1744] 5. Server: Based on the received question, the following response is generated: "The event that had the greatest impact on my career is..." and sent to the device. This is done in real time.

[1745] An example prompt might look like this:

[1746] "You are interviewing a well-known professional. The first question is, 'What has been the most influential event in your career so far?' Please continue with the following interview content."

[1747] Through this simulation, users can learn languages ​​through realistic conversational experiences, improving learning motivation and effectiveness.

[1748] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1749] Step 1:

[1750] Uploading conversation data

[1751] Terminal: The conversation data provider uses the terminal to select conversation data such as interview content or dialogue scenes, and clicks the upload button. The input is an audio file (e.g., .mp3) or a text file (e.g., .txt). The output is an upload request.

[1752] Step 2:

[1753] Receiving and storing data

[1754] Server: Receives conversation data sent from the device and temporarily stores it in storage. The input is the upload request from the device, and the output is the temporarily stored conversation data. Specifically, this includes checking the file size and format to ensure there are no inconsistencies.

[1755] Step 3:

[1756] Data Integrity Check

[1757] Server: Checks the integrity of the received data. The input is the temporarily saved conversation data, and the output is the status of whether it can be stored in the database. Specific operations include checking the completeness and consistency of the data.

[1758] Step 4:

[1759] Data storage and indexing

[1760] Server: Stores the integrity-checked data in a database and creates an index. The input is the checked conversation data, and the output is an indexed database entry, allowing for efficient searches.

[1761] Step 5:

[1762] Cleaning the data

[1763] Server: Cleans the conversation data stored in the database and removes unnecessary noise and privacy information. The input is the conversation data retrieved from the database, and the output is the clean data. Specifically, noise filtering is performed using a text mining library.

[1764] Step 6:

[1765] Creating a training set

[1766] Server: Converts the cleaned conversation data into a training set and feeds it to the LLM (large-scale language model). The input is clean data, and the output is the training set. Specifically, it normalizes and tokenizes the data.

[1767] Step 7:

[1768] Training a conversation model

[1769] Server: Trains the LLM using the training set. The input is the training set, and the output is the trained conversation model. Specific operations include updating the model parameters.

[1770] Step 8:

[1771] Saving the conversation model

[1772] Server: Stores the trained conversational model in a repository. The input is the trained conversational model, and the output is the model stored in the repository, so that it can be recalled and used later.

[1773] Step 9:

[1774] Accepting user requests

[1775] User: Logs in to the system through a terminal, selects a conversation partner and a situation, and requests a simulation. The input is the user's selection information, and the output is the request transmission.

[1776] Step 10:

[1777] Generate an initial response

[1778] Server: Based on the user's request, it calls the appropriate conversation model and generates an initial response. The input is the user's request, and the output is the generated initial response. The specific process is to input a prompt to the conversation model and get a response.

[1779] Step 11:

[1780] Viewing the response

[1781] Terminal: Displays the generated initial response on the screen. The input is the initial response from the server, and the output is the response displayed to the user.

[1782] Step 12:

[1783] Enter the next conversation phrase

[1784] User: Receives the initial response, enters the next conversational phrase, and sends it to the server through the terminal. The input is the next conversational phrase, and the output is a request to the server.

[1785] Step 13:

[1786] Producing the following response

[1787] Server: Analyzes the user's input, generates the next response, and sends it to the device. The input is the user's next conversation phrase, and the output is the generated response. This process is done in real time.

[1788] Step 14:

[1789] Save conversation logs

[1790] Server: Stores the interactions between the user and the conversation model as a log. The input is the data of each conversation turn, and the output is the saved conversation log. The log is also used as training data for subsequent tasks.

[1791] (Application example 1)

[1792] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1793] Today's language learners have limited opportunities to learn languages ​​efficiently through real-life conversation experiences with specific experts or celebrities. This issue can decrease learners' motivation and reduce learning effectiveness. Furthermore, content distribution services, in particular, lack platforms where users can enjoy interactive content. This leads to issues such as reduced user engagement and a decrease in the appeal of the service.

[1794] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1795] In this invention, the server includes: a means for a conversation data provider to upload conversation data from a terminal; a means for the server to store the received conversation data in a database; a means for the server to supply the stored conversation data to an LLM (large-scale language model) for training; a means for a user to select a conversation partner and situation from the terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the terminal; a means for the server to store the interaction between the user and the conversation model as a log; a means for the server to store the generated conversation log and provide it for later review or additional learning; and a means for the terminal to display responses to the user based on the selected conversation situation. This allows users to learn languages ​​through realistic conversation experiences with celebrities and experts, improving learning motivation and effectiveness. It also enables content distribution services to provide users with more interactive and engaging content.

[1796] A "conversation data provider" is an individual or organization that is responsible for uploading conversation data from a terminal.

[1797] "Terminal" refers to a device used by a conversation data provider or user to access the system, including a smartphone, tablet, head-mounted display, computer, etc.

[1798] "Server" is a computer system that receives, stores, and analyzes conversation data and manages and serves trained models.

[1799] The "database" is a system that centrally stores and manages received conversation data and generated conversation models.

[1800] An "LLM (Large-Scale Language Model)" is an artificial intelligence model trained on a huge amount of text data and capable of generating natural-sounding conversations.

[1801] A "repository" is a data storage system where trained conversational models are stored.

[1802] A "user" is an individual or group that accesses the system, selects a conversation partner and situation, and requests a simulation.

[1803] A "conversational model" is an AI model that is generated based on a trained LLM and is used to generate responses corresponding to specific conversational scenarios.

[1804] A "realistic conversational experience" is a process that simulates natural interactions that are close to real conversations.

[1805] A "conversation log" is a record of the interactions between a user and a conversation model, and is used as subsequent training data.

[1806] "Review" is the process of reviewing past conversation logs to enhance learning effectiveness.

[1807] The present invention provides a system for language learning that allows users to engage in realistic conversational simulations with celebrities and professionals, and is particularly suited to devices such as smartphones, tablets, and head-mounted displays.

[1808] System Configuration

[1809] 1. Hardware Configuration

[1810] Terminal: A device on which users can enjoy interactive conversation simulations. This includes smartphones, tablets, and head-mounted displays.

[1811] Server: A computer system that stores and manages conversation data, trains LLMs (large-scale language models), and generates responses.

[1812] Database: A system that centrally stores and manages conversation data and generated conversation models.

[1813] 2. Software Configuration

[1814] Flask: A microframework for building web applications that handles server-side processing.

[1815] OpenAI API: Used to generate natural-sounding conversational responses based on user input using generative AI models.

[1816] Data processing and calculation

[1817] Receiving user input: The server receives phrases entered by the user on the device, including text input and voice input.

[1818] Prompt Generation: Based on the user's input, generate a prompt to continue the virtual conversation, for example, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1819] Response Generation: Uses OpenAI's ChatGPT API to generate appropriate responses based on the generated prompts.

[1820] Returning the response: The generated response is sent to the terminal and displayed to the user.

[1821] Saving conversation logs: Records of interactions between the user and the conversation model (conversation logs) are saved on the server.

[1822] Specific examples

[1823] If a user is trying to learn Japanese, the process would be as follows:

[1824] 1. User: I want to simulate a conversation with a famous actor.

[1825] 2. User Input: "Tell me about your passion for filmmaking."

[1826] 3. Server: Takes input and generates the following prompt:

[1827] Prompt: "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1828] 4. OpenAI API: Generate appropriate responses based on prompts.

[1829] 5. Terminal: Displays the generated response to the user.

[1830] In this way, users can effectively learn languages ​​through realistic conversation experiences with celebrities and experts.

[1831] The above is the "Mode for Carrying Out the Invention."

[1832] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1833] Step 1:

[1834] The user selects a conversation partner and situation from the terminal and requests a simulation.

[1835] Input: User selection of conversation partner and situation.

[1836] Output: The request data to the server.

[1837] How it works: A user logs into the application using a smartphone, tablet, or head-mounted display, selects a conversation partner (e.g., a famous actor) and a situation for language learning, and then sends the selection to the server.

[1838] Step 2:

[1839] The server receives the request and invokes the conversation model.

[1840] Input: Request data from the user.

[1841] Output: The initial response of the conversation model.

[1842] Specific operation: Based on the received request, the server retrieves the corresponding conversation model from the database, for example, selects the conversation model of a famous actor, and generates an initial response.

[1843] Step 3:

[1844] The server receives the user's input and generates a prompt.

[1845] Input: A phrase entered by the user through the device (e.g., "Tell me about your passion for filmmaking.").

[1846] Output: The generated prompt statement.

[1847] Specific behavior: Based on the user's input phrase, the server generates a prompt sentence. The generated prompt sentence will be something like, "You are a famous actor. Please continue the following conversation:\nQ: Tell me about your passion for filmmaking.\nA:"

[1848] Step 4:

[1849] The server calls the OpenAI API to generate a response based on the prompt.

[1850] Input: The generated prompt statement.

[1851] Output: The generated response.

[1852] Specific operation: The server uses the OpenAI API to generate an appropriate response based on the generated prompt. At this stage, the prompt is sent to the API and the response text is retrieved.

[1853] Step 5:

[1854] The server generates a response that is sent to the terminal and displayed to the user.

[1855] Input: The response obtained from the OpenAI API.

[1856] Output: The response text that is displayed on the terminal.

[1857] Specific operation: The server sends the generated response to the terminal, and the user confirms this response through the terminal. For example, the response displayed is "About your passion for filmmaking."

[1858] Step 6:

[1859] The server stores the interaction between the user and the conversation model as a log.

[1860] Input: The dialogue between the user and the conversation model.

[1861] Output: Saved conversation logs.

[1862] How it works: The server stores the interactions between the user and the conversation model in a database for subsequent learning and review. Conversation logs are automatically saved and made available for the user to refer to later.

[1863] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1864] The present invention combines an emotion engine with a language learning system that allows users to learn languages ​​through realistic conversation simulations, generating optimal responses according to the user's emotional state. This system aims to improve learners' motivation and learning effectiveness by providing simulations based on conversation data with celebrities and professionals, in particular.

[1865] System Configuration

[1866] The system includes the following major components:

[1867] 1. Terminal

[1868] The device (e.g., computer, smartphone, tablet) through which the conversation data provider or user accesses the system.

[1869] An interface for conversation data providers to upload conversation data.

[1870] An interface for users to request conversation simulations and view the simulation results.

[1871] 2. Server

[1872] The ability to receive, save, clean, and store conversation data in a database.

[1873] A function that trains LLMs (large-scale language models) based on saved conversation data.

[1874] Ability to save trained conversational models and recall them whenever needed.

[1875] The ability to use conversational models to generate realistic responses based on user requests.

[1876] Ability to recognize user emotions using an emotion engine and generate optimal responses.

[1877] A function that saves conversation logs and uses them as training data for future use.

[1878] 3. Database

[1879] A system that centrally stores and manages received conversation data and generated conversation models.

[1880] 4. Emotion Engine

[1881] An engine for recognizing emotions from user text input and voice data and optimizing responses based on that.

[1882] Collection and handling of conversation data

[1883] Terminal: Conversation data providers use a terminal to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. At this time, audio files and text files can be used.

[1884] Server: Uploaded conversation data is received by the server and temporarily stored in storage. The server then checks the integrity of the received data and stores it in a database. The stored data can be efficiently managed and searched by creating an index.

[1885] Training the model

[1886] Server: The stored conversation data undergoes a data cleaning process to remove noise and privacy information. The cleaned data is converted into a training set and fed to the LLM, which generates a trained conversation model that learns the speaking style and phrases of a specific person.

[1887] User Interaction

[1888] User: Logs in to the system, checks his / her profile information, and then selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format).

[1889] Terminal: The selection is sent to the server as a request.

[1890] Server: Accepts requests, invokes the appropriate conversation model, and generates an initial response that is sent to the device and displayed to the user.

[1891] Real-time conversation simulation and emotion recognition

[1892] Emotion engine: Recognizes emotions from the user's text input or voice data. The recognized emotion information is sent to the server.

[1893] Server: Optimizes responses based on the emotional information received from the emotion engine. Sends the generated responses to the device so that the conversation continues in real time.

[1894] User: Enter the following conversation phrase and send it to the server via the device. If an emotion is recognized, that information is also sent.

[1895] Server: Analyzes the user's input data and emotional information, and generates the following response, which is also sent to the device in an optimized form.

[1896] Server: These conversation logs are stored and used as training data for future use, improving the accuracy and quality of the system.

[1897] Specific examples

[1898] A user simulates an interview with celebrity B.

[1899] 1. Terminal: The user logs in to the system and selects celebrity B as the conversation partner. The situation is chosen as an interview format.

[1900] 2. Server: Based on the user request, it invokes Celebrity B's conversation model and generates an initial response.

[1901] 3. Terminal: The initial response is displayed to the user.

[1902] 4. User: Enter the following question and send it to the server via the terminal.

[1903] 5. Emotion engine: Recognizes emotions from the user's input data and also sends this information to the server.

[1904] 6. Server: Based on the emotion information received from the emotion engine, it generates and optimizes the next response and sends it to the terminal.

[1905] 7. Server: Saves the conversation log.

[1906] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[1907] The processing flow will be explained below.

[1908] Step 1: Upload your conversation data

[1909] Terminal: Conversation data providers use terminals to upload conversation data (voice or text) to the system.

[1910] Server: Receives conversation data and temporarily stores it in storage.

[1911] Step 2: Verify the integrity of the conversation data and store it

[1912] Server: Checks the integrity of the received conversation data and stores it in a database, adding appropriate metadata when storing it.

[1913] Step 3: Indexing

[1914] Server: Indexes the stored conversation data, enabling efficient search and management.

[1915] Step 4: Data Cleaning

[1916] Server: Cleans the stored conversation data to remove noise and privacy information.

[1917] Step 5: Generate a training set

[1918] Server: Converts the cleaned data into a training set and feeds it into the LLM (large-scale language model).

[1919] Step 6: Train and save the model

[1920] Server: Trains the LLM using the training set to learn the speaking style and phrases of a particular person.

[1921] Server: Stores the trained conversation model in a repository.

[1922] Step 7: Logging in users

[1923] User: Log in to the system and check your profile information.

[1924] Terminal: Sends login information to the server for authentication.

[1925] Step 8: Choose your conversation partner and situation

[1926] User: Selects a conversation partner (e.g., celebrity A) and a situation (e.g., interview format) from the system.

[1927] Terminal: Sends the selection to the server.

[1928] Step 9: Receiving a request and generating an initial response

[1929] Server: Receives requests from users, invokes the selected conversation model, and generates the initial response.

[1930] Server: Sends the generated initial response to the terminal.

[1931] Step 10: View the initial response

[1932] Terminal: Displays the initial response received from the server to the user.

[1933] Step 11: Enter conversation phrases

[1934] User: Enter the following conversation phrase into your device:

[1935] Terminal: Sends the entered data to the server.

[1936] Step 12: Recognize emotions

[1937] Server: Uses an emotion engine to recognize emotions from user input data and voice data.

[1938] Step 13: Generate and optimize responses

[1939] Server: Generates the optimal response based on the user's input data and recognized emotional information.

[1940] Server: Sends the generated optimization response to the terminal.

[1941] Step 14: View the response

[1942] Terminal: displays the optimization response received from the server to the user.

[1943] Step 15: Save the conversation log

[1944] Server: Saves conversation logs with users in a database and uses them as training data for future use.

[1945] Specific examples

[1946] A user simulates an interview with celebrity B.

[1947] 1. Step 1:

[1948] Device: Celebrity B uploads his / her interview recording.

[1949] Server: Receives and temporarily stores conversation data.

[1950] 2. Step 2:

[1951] Server: Checks the integrity of the data and stores it in the database.

[1952] 3. Step 3:

[1953] Server: Creates indexes to simplify later searches.

[1954] 4. Steps 4-6:

[1955] Server: Performs data cleaning, generates training sets, and trains the LLM.

[1956] 5. Steps 7-8:

[1957] User: Log in to the system and select Celebrity B. Select the interview format.

[1958] Terminal: Sends the request contents to the server.

[1959] 6. Steps 9-10:

[1960] Server: Generates an initial response and sends it to the terminal.

[1961] Terminal: Display the initial response to the user.

[1962] 7. Steps 11-12:

[1963] User: Enter the following question and send it to the server via the terminal.

[1964] Server: Recognizes emotions using an emotion engine.

[1965] 8. Steps 13-14:

[1966] Server: Generates the optimal response based on the recognized emotional information and sends it to the device.

[1967] Terminal: Display the optimization response to the user.

[1968] 9. Step 15:

[1969] Server: Saves conversation logs and uses them as training data for future use.

[1970] In this way, users can advance their language learning while simulating a realistic and emotionally relevant interview experience with Celebrity B.

[1971] Example 2

[1972] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1973] Existing language learning systems have difficulty enabling users to effectively learn a language through realistic conversation simulations. Furthermore, they lack a mechanism for generating optimal responses based on the user's emotional state. As a result, learners' motivation and learning effectiveness may decrease.

[1974] The specification processing by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a conversation data provider to upload conversation data from an information terminal; a means for the server to save the received conversation data in a database; a means for the server to supply the saved conversation data to a large-scale language model and train it; a means for the server to save the trained conversation model in a repository; a means for a user to select a conversation partner and a situation from an information terminal and request a simulation; a means for the server to call the conversation model in response to the user's request, generate a response, and send it to the information terminal; a means for an emotion engine to recognize emotions from the user's input data and send the emotion information to the server; a means for the server to optimize and generate a response based on the emotion information and send it to the information terminal; and a means for the server to save the interaction between the user and the conversation model as a log. This enables users to effectively learn language through realistic conversation simulations and improve their learning experience.

[1975] "Conversation Data Provider" means a person or organization that provides conversation data to the system.

[1976] An "information terminal" is an electronic device used to access a system and send and receive data. Examples include computers, smartphones, and tablets.

[1977] A "server" is a computer system responsible for receiving, processing, and storing conversation data.

[1978] "Conversation Data" refers to a record of a conversation provided as an audio file or text file.

[1979] A "database" is a storage system that allows the server to efficiently store, manage, and search conversation data received.

[1980] A "large-scale language model" is an algorithmic model for natural language processing that is trained on a large amount of text data.

[1981] "Training" refers to the process of feeding conversational data into a large-scale language model to learn specific speaking styles and phrases.

[1982] A "repository" is a storage system for storing trained conversational models.

[1983] "Profile information" refers to personal information and setting data that a user confirms after logging in to the system.

[1984] A "situation" is a specific situation or context in a conversation simulation, such as an interview format.

[1985] "Simulation" is the process of engaging in virtual interactions based on user-selected conversation partners and situations.

[1986] An "emotion engine" is an algorithm or system that recognizes emotions from user input data and provides the results to a server.

[1987] "Input data" refers to text and voice data provided by a user to a system.

[1988] "Emotional information" refers to data that indicates the emotional state of the user as recognized by the emotion engine.

[1989] "Log" refers to historical data that records interactions between users and conversation models.

[1990] "Optimization" refers to the process of appropriately adjusting the content and expression of a response based on recognized emotional information.

[1991] MODE FOR CARRYING OUT THE INVENTION

[1992] The present invention combines a system in which users learn languages ​​through realistic conversation simulations with an emotion engine to generate optimal responses according to the user's emotional state. The system includes the following main components:

[1993] 1. Terminal

[1994] A terminal is a device through which a data provider or user accesses the system. Examples include computers, smartphones, and tablets. Terminals are equipped with the following interfaces:

[1995] An interface for conversation data providers to upload conversation data

[1996] An interface for users to request conversational simulations and view the results of the simulations

[1997] 2. Server

[1998] The server receives, stores, cleans, and stores conversation data in a database. The server has the following functions:

[1999] 1. A function to receive conversation data and temporarily store it in storage

[2000] 2. A function to check the integrity of received data and store it in a database

[2001] 3. Ability to clean stored data and remove noise and private information

[2002] 4. The ability to convert cleaned data into a training set and feed it into a large-scale language model (LLM).

[2003] 5. Ability to save the conversation model generated through training and recall it whenever needed

[2004] 6. Ability to use conversational models to generate realistic responses based on user requests

[2005] 7. Ability to recognize user emotions using an emotion engine and generate optimal responses

[2006] 8. Ability to save conversation logs and use them as training data for future use

[2007] 3. Database

[2008] The database is a system that centrally stores and manages received conversation data and generated conversation models. Indexing enables efficient management and search.

[2009] 4. Emotion Engine

[2010] The emotion engine is an engine that recognizes emotions from user text input and voice data and optimizes responses based on those emotions.

[2011] Collection and processing of conversation data

[2012] Users use their devices to upload their own conversation data (e.g., interview content, dialogue scenes) to the system. Audio files and text files can be used for this purpose. The device then sends the conversation data uploaded by the user to the server. The server checks the integrity of the received data and stores it in a database. Indexing allows for efficient management and searching.

[2013] Training the model

[2014] The server retrieves conversation data from the database and performs a data cleaning process. It converts the cleaned data into a training set and supplies it to the LLM. This generates a trained conversation model that has learned the speaking style and phrases of a specific person. The server stores the generated conversation model in a repository and can be called up whenever needed.

[2015] Preparing for user interaction

[2016] A user logs in to the system using a terminal, checks their profile information, and selects a conversation partner and situation. The terminal then sends the selection to the server as a request. The server accepts the request, selects an appropriate conversation model from its saved storage, generates an initial response, and sends it to the terminal.

[2017] Real-time conversation simulation and emotion recognition

[2018] The user inputs the next question or dialogue phrase and sends it to the server via the device. The emotion engine analyzes the user's input data and recognizes emotions. The recognized emotion information is sent to the server. The server then optimizes the response based on the emotion information and sends the generated response to the device. This allows the user to effectively learn the language through realistic conversation simulation. The server saves these conversation logs and uses them as subsequent training data to improve the accuracy and quality of the system.

[2019] Specific examples

[2020] A user simulates an interview with a celebrity

[2021] 1. Terminal: The user logs in to the system and selects a celebrity to talk to. The situation is chosen to be an interview format.

[2022] 2. Server: Based on the user request, the celebrity conversation model is called and an initial response is generated, which reads, "Hello, celebrity. What would you like to talk about today?"

[2023] 3. Terminal: The initial response is displayed to the user. Type "Hello, celebrity. I'd like to ask you my next question. Tell me about your latest project."

[2024] 4. Emotion engine: Recognizes emotions from the user's input data and sends that information to the server. It is analyzed as "curious expression."

[2025] 5. Server: Based on the emotion information received from the emotion engine, the server generates and optimizes a response and sends it to the device. The response generated is "Thank you. Now, I'll tell you about my latest project..." and is displayed to the user.

[2026] 6. Server: Stores user interaction logs and uses them as training data for subsequent use to improve the accuracy of the system.

[2027] This allows users to advance their language learning while enjoying a realistic interview experience.

[2028] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2029] Specific processing steps of the system

[2030] Step 1:

[2031] A user logs into a system using an information terminal. The input is the user's authentication information (e.g., username, password), and the output is the start of a session for the authenticated user. To access the system, the user enters the appropriate authentication information. The system receives this, authenticates, and then starts the user session.

[2032] Step 2:

[2033] Users upload their own conversation data from their information terminals. The input is an audio file or a text file, and the output is the conversation data sent to the server. Using the system's interface, conversation data providers select a data file and upload it to the server.

[2034] Step 3:

[2035] The server temporarily stores the received conversation data in storage. The input is the conversation data sent from the terminal, and the output is the data temporarily stored in storage. After receiving the data, the server verifies it, performs an integrity check, and then stores the data in storage.

[2036] Step 4:

[2037] The server stores the saved conversation data in a database. The input is the data stored in temporary storage, and the output is the data indexed in the database. When storing the data in the database, an index is created to enable efficient management and search.

[2038] Step 5:

[2039] The server retrieves the conversation data from the database and runs a data cleaning process. The input is the raw data in the database, and the output is the cleaned data. The dataset is cleaned using algorithms to remove noise and privacy information.

[2040] Step 6:

[2041] The server converts the cleaned conversation data into a training set and feeds it to an LLM (large-scale language model). The input is the cleaned data, and the output is the training set. This completes the dataset for training the conversation model.

[2042] Step 7:

[2043] The server saves the conversation model generated by training in a repository. The input is the conversation model generated by training, and the output is the model saved in the repository. The generated model is saved in storage so that it can be used in response to any request.

[2044] Step 8:

[2045] The user selects a conversation partner and a situation using an information terminal. The input is the selection information of the conversation partner and the situation, and the output is request data to the server. The details of the conversation simulation are specified through the interface, and the information is sent to the system.

[2046] Step 9:

[2047] The server calls the appropriate stored conversation model based on the user's request and generates an initial response. The input is the user's request data, and the output is the generated initial response. The specified model is loaded and an initial response according to the request is generated.

[2048] Step 10:

[2049] The server sends the generated initial response to the information terminal. The input is the generated initial response, and the output is the response displayed on the terminal. The generated response is forwarded to the user for display.

[2050] Step 11:

[2051] The user inputs the next question or dialogue phrase and sends it to the server through the terminal. The input is the user's next question or dialogue phrase, and the output is the data sent to the server. The user inputs the next step to continue the dialogue.

[2052] Step 12:

[2053] The emotion engine analyzes the user's input data and recognizes emotions. The input is the user's next question or dialogue phrase, and the output is the recognized emotion information. It analyzes the provided data and recognizes the user's emotional state.

[2054] Step 13:

[2055] The emotion engine sends the recognized emotion information to the server. The input is the recognized emotion information, and the output is the data sent to the server. The analysis results are transferred to the server.

[2056] Step 14:

[2057] The server optimizes and generates a response based on the emotional information. The input is the emotional information and the user's next question, and the output is an optimized response. The response content is adjusted and generated according to the recognized emotion.

[2058] Step 15:

[2059] The server sends the generated response to the terminal and displays it to the user. The input is the generated response and the output is the response displayed on the terminal. The server forwards the response to the user and continues the dialogue.

[2060] Step 16:

[2061] The server stores the interaction between the user and the conversational model as a log. The input is the interaction data between the user and the conversational model, and the output is the stored log. The stored log is used for subsequent training.

[2062] (Application example 2)

[2063] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2064] Conventional language learning systems have had difficulty providing optimal responses based on conversational realism and emotions. It has also been difficult for users to engage in natural conversations in virtual stores while selecting products and receiving information in real time. This has led to problems such as reduced learning effectiveness and user satisfaction.

[2065] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2066] In this invention, the server includes means for analyzing the emotional state of the user using an emotion engine and optimizing a response based on that information, means for providing the optimized response to the user's terminal in real time, and means for the user to select products and provide information in the virtual environment using the terminal, thereby enabling the user to obtain product information while engaging in natural conversation according to their emotions in real time in the virtual store.

[2067] A "conversation data provider" is a person or organization that uploads conversation data from a terminal.

[2068] "Terminal" means a device (e.g., computer, smartphone, tablet) used by a conversation data provider or user to access the system.

[2069] The "server" is the core computer system that receives and stores conversation data, serves it as training data, and ultimately generates responses.

[2070] A "database" is a system for storing and managing received conversation data and training data.

[2071] An "LLM (Large Scale Language Model)" is a language model trained using a large number of datasets to achieve natural language processing.

[2072] "Training" is the process of optimizing a large-scale language model using stored conversational data.

[2073] A "repository" is a system for storing and managing trained conversation models.

[2074] "User" means a person or entity that uses a terminal to access the system and use the conversation simulation.

[2075] An "emotion engine" is a technology that analyzes a user's emotional state and optimizes responses based on that information.

[2076] A "virtual environment" is a simulated environment that allows users to have a realistic experience within a digital space.

[2077] "Product selection" is the action of a user selecting a product of interest within a virtual environment.

[2078] "Information provision" means providing detailed information about the product selected by the user through a virtual assistant or the like.

[2079] "Real-time" means that processing is done immediately, without delay, and the results are provided to the user.

[2080] The following describes in detail an embodiment of the invention in a virtual store customer service system. First, the process of generating a program for the system that realizes this application example and its specific operation will be described.

[2081] Overall system configuration

[2082] The system consists of the following main elements:

[2083] 1. Terminal: A device through which a user accesses the system, such as a smartphone or head-mounted display.

[2084] 2. Server: The core computer system that receives, stores, trains, and generates responses from conversation data.

[2085] 3. Database: A system that stores and manages conversation data and conversation models.

[2086] 4. Emotion Engine: Technology that analyzes the user's emotional state and optimizes responses based on that information.

[2087] Detailed Operation

[2088] 1. User visits the virtual store:

[2089] By launching the application using a device (smartphone or head-mounted display), the user enters the 3D environment of the virtual store.

[2090] 2. Product Selection and Information:

[2091] Once the user selects the product of interest, product information is provided by the virtual assistant (trained conversation model).

[2092] 3. Real-time conversation simulation:

[2093] When a user asks a question to a virtual assistant, the emotion engine analyzes the question and the user's emotional state. The emotion engine's API (e.g., IBM Watson Tone Analyzer, Microsoft Azure Emotion API) analyzes the user's voice data and text to recognize the emotional state.

[2094] 4. Optimal response generation:

[2095] Based on the recognized emotional state, a large-scale language model (e.g., OpenAI's GPT series) generates an optimal response. The server sends a request to the triggered model and obtains the inference result.

[2096] 5. Response display:

[2097] The generated responses are displayed to the user through the device, and the conversation logs are saved by the server and used as training data for future use.

[2098] Specific hardware and software

[2099] Hardware

[2100] Smartphones: iPhone, Android, etc.

[2101] Head-mounted displays: Oculus Quest 2, HoloLens 2, etc.

[2102] software

[2103] Front-end frameworks: Unity, React Native, etc.

[2104] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.

[2105] Large-scale language models: OpenAI GPT-3, GPT-4, etc.

[2106] Database: MongoDB, MySQL, etc.

[2107] Server-side frameworks: Node.js, Django, etc.

[2108] Examples of concrete examples and prompts

[2109] Here's how the system works:

[2110] The user puts on a head-mounted display, enters a virtual store, and approaches specific products.

[2111] Example prompt: "What features does this new smartwatch have?"

[2112] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[2113] This allows users to obtain product information while engaging in natural, emotionally-responsive dialogue in real time in a virtual store.

[2114] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2115] Step 1:

[2116] Access to the system and entry to the virtual store

[2117] The user launches the application using a smartphone or head-mounted display, which accesses the 3D environment of the virtual store, where the user logs in to the system and the initial interface is displayed.

[2118] Input: User login information, device type

[2119] Data processing / calculation: User authentication, display of virtual store using 3D rendering engine

[2120] Output: Login success message, virtual store UI display

[2121] Step 2:

[2122] Product selection and initial information provision

[2123] The user moves through the virtual environment and selects a specific product. The device sends this selection information to the server, which retrieves the product's initial information from a database and displays it.

[2124] Input: ID of the product selected by the user

[2125] Data processing / calculation: database query for product information, information format conversion

[2126] Output: Product details (text, images, etc.)

[2127] Step 3:

[2128] Start the conversation simulation

[2129] The user inputs a question related to the product, and the device sends the question to the server, which then uses an emotion engine to analyze the user's emotional state.

[2130] Input: User text input (question content)

[2131] Data processing / calculation: Emotion analysis using emotion engines (e.g., IBM Watson Tone Analyzer)

[2132] Output: Emotional state data (e.g., surprise, joy, dissatisfaction)

[2133] Step 4:

[2134] Generating the best response

[2135] The server uses the emotional state data and the user's question to trigger a large-scale language model (e.g., OpenAI GPT-4) to generate an optimal response.

[2136] Input: User question, emotional state data

[2137] Data processing / calculation: Prompt generation and response generation for large-scale language models

[2138] Output: Optimized response text

[2139] Step 5:

[2140] Viewing the response

[2141] The generated response text is sent from the server to the terminal, which displays the response to the user.

[2142] Input: Response text data from the server

[2143] Data processing / calculation: rendering text data, updating the interface

[2144] Output: User confirms response text

[2145] Step 6:

[2146] Save conversation logs

[2147] The conversational exchanges are logged and stored by the server and used as training data for subsequent use.

[2148] Input: Conversation data between the user and the virtual assistant

[2149] Data processing / calculation: Logging to database, indexing

[2150] Output: Saved conversation log

[2151] Examples:

[2152] Prompt Sentence Examples

[2153] A user asks: "What features does this new smartwatch have?"

[2154] The virtual assistant responds: "This smartwatch has many features, including heart rate monitoring, GPS, music playback, and message notifications. Would you like to give it a try?"

[2155] This step allows users to receive appropriate responses in real time based on their emotions, enabling an efficient and satisfying virtual store experience.

[2156] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2157] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2158] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2159] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2160] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2161] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2162] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2163] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2164] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2165] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2166] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2167] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2168] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2169] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2170] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2171] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2172] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2173] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2174] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2175] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2176] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2177] The following is further disclosed regarding the above embodiment.

[2178] (Claim 1)

[2179] A means for a conversation data provider to upload conversation data from a terminal;

[2180] A means for storing the conversation data received by the server in a database;

[2181] A means for the server to feed the stored conversation data to a large-scale language model (LLM) for training;

[2182] A means for the server to store the trained conversational model in a repository; and

[2183] A means for users to select a conversation partner and situation from their device and request a simulation;

[2184] A means for the server to invoke the conversation model in response to a user request, generate a response, and send it to the terminal;

[2185] The system includes a means for the server to store a log of interactions between the user and the conversation model.

[2186] (Claim 2)

[2187] 10. The system of claim 1, wherein the server further comprises means for cleaning the stored conversation data to remove unwanted noise and privacy information.

[2188] (Claim 3)

[2189] 10. The system of claim 1, further comprising means for the server to convert the cleaned conversation data into a training set and provide it to the LLM.

[2190] "Example 1"

[2191] (Claim 1)

[2192] A means for a conversation data provider to upload conversation data from a terminal;

[2193] A means for storing the conversation data received by the server in a database;

[2194] A means for the server to feed the stored conversation data to a large-scale language model (LLM) for training;

[2195] A means for the server to store the trained conversational model in a repository; and

[2196] A means for users to select a conversation partner and situation from their device and request a simulation;

[2197] A means for the server to invoke the conversation model in response to a user request, generate a response, and send it to the terminal;

[2198] a real-time conversation simulation means for allowing a user to input a next conversation phrase and for a server to generate a response and transmit it to the terminal;

[2199] The system includes a means for the server to store a log of interactions between the user and the conversation model.

[2200] (Claim 2)

[2201] 10. The system of claim 1, wherein the server further comprises means for cleaning the stored conversation data to remove unwanted noise and privacy information.

[2202] (Claim 3)

[2203] 10. The system of claim 1, further comprising means for the server to convert the cleaned conversation data into a training set and provide it to the LLM.

[2204] "Application Example 1"

[2205] (Claim 1)

[2206] A means for a conversation data provider to upload conversation data from a terminal;

[2207] A means for storing the conversation data received by the server in a database;

[2208] A means for the server to feed the stored conversation data to a large-scale language model (LLM) for training;

[2209] A means for the server to store the trained conversational mod...

Claims

1. A means for a conversation data provider to upload conversation data from a terminal; A means for storing the conversation data received by the server in a database; A means for the server to feed the stored conversation data to a large-scale language model (LLM) for training; A means for the server to store the trained conversational model in a repository; and A means for users to select a conversation partner and situation from their device and request a simulation; A means for the server to invoke the conversation model in response to a user request, generate a response, and send it to the terminal; The system includes a means for the server to store a log of interactions between the user and the conversation model.

2. 10. The system of claim 1, wherein the server further comprises means for cleaning the stored conversation data to remove unwanted noise and privacy information.

3. The system of claim 1, further comprising means for the server to convert the cleaned conversation data into a training set and provide it to the LLM.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A